Pith. sign in

Paper Citation Record · LEDGER

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One

As of 18 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2505.15306.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15306 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:25:00.054246Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7ad5cac9-df3a-44a8-bfda-9bd22a7c4f8b · outbound

This paper cites Reinforcement learning: An introduction.A Bradford Book, 2018.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement learning: An introduction.A Bradford Book, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.681461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:56.234369Z digest=sha256:bc010346c8b02280e1c8a44659ec1f073315333e1eb2eaca0e255371505c6236

Observation b65d7d72-dd7a-4714-becf-e4a3fe7b93d8 · outbound

This paper cites An introduction to deep reinforcement learning.Foundations and Trends® in Machine Learning, 11(3-4):219–354, 2018.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One An introduction to deep reinforcement learning.Foundations and Trends® in Machine Learning, 11(3-4):219–354, 2018

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.660069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:56.288962Z digest=sha256:0bd7277a974af6b1794522faf5bca22c34c7bc737a1f8d744ad655e1fd94ad7f

Observation e379d43f-4d11-4265-8313-801cd0839fad · outbound

This paper cites Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Mastering the game of go without human knowledge.nature, 550(7676):354–359, 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.346937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.346937Z digest=sha256:d487cb3b8453a2151cb1674731e8166cafc6f10aaaa4456512d4b9b91f2df4e4

Observation 2db00582-0d4f-4bac-856b-d4dcff09eaab · outbound

This paper cites Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Grandmaster level in starcraft ii using multi-agent reinforcement learning.nature, 575(7782):350–354, 2019

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.431165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.431165Z digest=sha256:04ef32292e6c78ad6ca775878505aa2483b12135f08e048725ffac2e6eb76266

Observation b5c85c06-0d0d-4240-a7f2-d61777df09e3 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Dota 2 with Large Scale Deep Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:56.541701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:56.541701Z digest=sha256:89cf65c7ef5aedd1c486da6c274d9de5a62ffcf4ba2bc3f5146a1ea0b5e5e387

Observation eb94c72c-b50b-45d6-98d8-4ab6433b0bf5 · outbound

This paper cites Mastering atari games with limited data.Advances in neural information processing systems, 34:25476–25488, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Mastering atari games with limited data.Advances in neural information processing systems, 34:25476–25488, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.606430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:56.598174Z digest=sha256:d5d7e84a04cfa108a6a1acf44788041503b098cc2f4ff3a5e2b5a2db52d28010

Observation 78b34cf3-9b9d-4b81-85d1-48c022f7a2b8 · outbound

This paper cites A graph placement methodology for fast chip design.Nature, 594(7862):207–212, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A graph placement methodology for fast chip design.Nature, 594(7862):207–212, 2021

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.584564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:56.693310Z digest=sha256:b39fbbf3d116fcb99497a523e3c7659a76c5acfaf74146f406a105ac0df32b79

Observation d14d986c-aa0b-4a88-a7b6-4179af0f8d06 · outbound

This paper cites Hierarchical reinforcement learning for scarce medical resource allocation with imperfect information.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Hierarchical reinforcement learning for scarce medical resource allocation with imperfect information

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.558171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:56.767959Z digest=sha256:9c82a2184d82cb4a2c43fdbe276ab0b5aee5edeb3e4791f6a887d995e84f6ea7

Observation 99e60500-87cd-4591-81ca-4282dfb822a5 · outbound

This paper cites Reinforcement learning enhances the experts: Large-scale covid-19 vaccine allocation with multi-factor contact network.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement learning enhances the experts: Large-scale covid-19 vaccine allocation with multi-factor contact network

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.538975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:56.855259Z digest=sha256:6b158f660d7591d678a4dbbf2834faef11a4ff104db1b3a90b52e8fdb6c289db

Observation f0e78dcd-1a14-4ab8-9db4-436492db1d43 · outbound

This paper cites Gat-mf: Graph attention mean field for very large scale multi-agent reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Gat-mf: Graph attention mean field for very large scale multi-agent reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.520522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:56.945786Z digest=sha256:9296f154b9b401d5574ce7ce77737537b85a8a17f7ffcb6b10db2d2d382ef8e5

Observation 1638ec47-1810-47be-9ac0-132adbf5f3bb · outbound

This paper cites Spatial planning of urban communities via deep reinforcement learning.Nature Computational Science, 3(9):748– 762, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Spatial planning of urban communities via deep reinforcement learning.Nature Computational Science, 3(9):748– 762, 2023

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.047385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.047385Z digest=sha256:a8c98bc061f9d24af21666c586f24de119274ee6d50a6c6b464ca9e70de4dd84

Observation 15e054f8-c016-49d9-a301-117d0711c46d · outbound

This paper cites A survey of machine learning for urban decision making: Applications in planning, transportation, and healthcare.ACM Computing Surveys, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey of machine learning for urban decision making: Applications in planning, transportation, and healthcare.ACM Computing Surveys, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.491251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:57.099196Z digest=sha256:984025822c7f7689f3966ce17ef4c88bc6f3fc8999149ba5e30624ce43b1da63

Observation 39339997-605b-4318-bbbf-bfec067c9de8 · outbound

This paper cites Dyps: Dynamic parameter sharing in multi-agent reinforcement learning for spatio-temporal resource allocation.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Dyps: Dynamic parameter sharing in multi-agent reinforcement learning for spatio-temporal resource allocation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.471478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:57.212195Z digest=sha256:51bc3447741d2c480f9d295c217e86718b9f874981425bf5f2bc9256a4cfed76

Observation 3652024f-61f0-4ba1-b24b-bf63788475f8 · outbound

This paper cites Coopride: Cooperate all grids in city-scale ride-hailing dispatching with multi-agent reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Coopride: Cooperate all grids in city-scale ride-hailing dispatching with multi-agent reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.453447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:57.301068Z digest=sha256:2137be3bf57fb85ce5a8fecf7cda07ae01f81f8892d5b877b6196bfa8fa98adf

Observation a86bb074-2a9b-47c2-90ba-95720d7c2bec · outbound

This paper cites A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.402374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.402374Z digest=sha256:712d1b90e092478476320703635f52dc22fe165f46a5bae01c1bfe9fb64fc969

Observation 52045072-bd79-4cbd-a6a1-4dafc2302c9e · outbound

This paper cites How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.516895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.516895Z digest=sha256:b5337bad42861d9b62a7c58face79807285c12d21aee33c6ae85a316521ad55b

Observation 73c5c68b-f8d5-41cd-840a-450a3b80d005 · outbound

This paper cites Ensemble deep learning: A review.Engineering Applications of Artificial Intelligence, 115:105151, 2022.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble deep learning: A review.Engineering Applications of Artificial Intelligence, 115:105151, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.437638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:57.614737Z digest=sha256:a1a1cc253408b9b6fdacb8e94875c9918bcd985f36e9c5bd949b4404584f45fb

Observation 85e2220f-f902-4ad9-a9af-b91f988f89e4 · outbound

This paper cites A survey on ensemble learning.Frontiers of Computer Science, 14:241–258, 2020.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey on ensemble learning.Frontiers of Computer Science, 14:241–258, 2020

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.421082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:57.701960Z digest=sha256:b6fa31960f6e3a85783b2f33e1b5c55146c0a6c56364e8ca42590620a0b52360

Observation 422ef0c2-f114-4985-bf1d-8e18ca44d06a · outbound

This paper cites Ensemble reinforcement learning: A survey.Applied Soft Computing, page 110975, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble reinforcement learning: A survey.Applied Soft Computing, page 110975, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.405209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:57.775550Z digest=sha256:63aabd9a959e2e4c0d1227173296a17704ec0c12469b6e148f2e3456078aa9ed

Observation 69160df6-5acd-4af3-b01f-1599a827a21e · outbound

This paper cites Neural network ensembles in reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Neural network ensembles in reinforcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.389623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:57.856542Z digest=sha256:6645a0ad82a8a338fc93626fc5449c0018949b55d576f14011f3fea99bb17feb

Observation 499cc08f-ef3e-4042-b788-fac959861c37 · outbound

This paper cites Reinforcement Learning based dynamic weighing of Ensemble Models for Time Series Forecasting.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Reinforcement Learning based dynamic weighing of Ensemble Models for Time Series Forecasting

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.963590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.963590Z digest=sha256:a377abf2dcce58f8207830de50215b0ddedd4ada29bc573c48cbc2a284a810d6

Observation f5a597ba-0a2c-4a3b-b84f-36281b50f99a · outbound

This paper cites Ensemble algorithms in reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38(4):930–936, 2008.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Ensemble algorithms in reinforcement learning.IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 38(4):930–936, 2008

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.374081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:58.083139Z digest=sha256:46af996ebb3507897632ead76f9f63a8dcc3bc364abca84bbced8690d6d48df8

Observation cc9b4d82-1585-4b67-a68d-64dd57681529 · outbound

This paper cites The arcade learning environment: An evaluation platform for general agents.Journal of Artificial Intelligence Research, 47:253–279, 2013.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One The arcade learning environment: An evaluation platform for general agents.Journal of Artificial Intelligence Research, 47:253–279, 2013

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.358572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:58.191168Z digest=sha256:565ff58f02dddedd227a30562d22db10c2b539a413cead85f444569a8d0e16de

Observation 9e8f1398-449e-4f9f-a24d-73586e8337f3 · outbound

This paper cites Model-Based Reinforcement Learning for Atari.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Model-Based Reinforcement Learning for Atari

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.285479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.285479Z digest=sha256:a72d00067c8be58de5d693000d661bbf71bef9b462d2c48a757031591a18e598

Observation 80b1eed5-579f-40ec-9f56-ebcb17ade291 · outbound

This paper cites an unresolved cited work.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.360163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.360163Z digest=sha256:349ab0a11b67832bd9e0ed467f6a15cd69cb7eeee4509398bfb3eb0ebf2beb1f

Observation e4ac7fd4-1ae1-47eb-bc30-7a4d805c4141 · outbound

This paper cites A survey of gpt-3 family large language models including chatgpt and gpt-4.Natural Language Processing Journal, page 100048, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey of gpt-3 family large language models including chatgpt and gpt-4.Natural Language Processing Journal, page 100048, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.330048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:58.423192Z digest=sha256:96252bc052d69c05d2144ef9943b8d868d8c432107bd822bbf82913a0942af88

Observation 2cec5757-1406-40ba-8e63-a0a23a7a4af5 · outbound

This paper cites GPT-4 Technical Report.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.507475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.507475Z digest=sha256:dcdab0fdde807adcf286e4cd0fbd681c5fde1baf302bddc4d70bed363f021163

Observation 9ed98770-0f4d-4d91-a149-51d27d2c74ac · outbound

This paper cites Evaluation of openai o1: Opportunities and challenges of agi.arXiv preprint arXiv:2409.18486, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Evaluation of openai o1: Opportunities and challenges of agi.arXiv preprint arXiv:2409.18486, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.596746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.596746Z digest=sha256:9f20dbdefca96354342ca7586189459dec88fd7c7440cfb3b845dc680c8da403

Observation 0bd3492f-c5fb-4285-91ba-dc3f233a93e9 · outbound

This paper cites Early access for safety testing.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Early access for safety testing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.719213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.719213Z digest=sha256:01fc164e38f54a8f7a049a670499229298daf827efddaafb140ce043dc3c96dc

Observation 7e419586-5f54-4b2f-97a6-d767d7a28875 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.773436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.773436Z digest=sha256:2ceaa7216e0b93ed39cd447f360ef816777c06d0cd0b95cff82265da489915e8

Observation 81c4a7cd-d620-47d9-aa37-efa2acf441b9 · outbound

This paper cites The Llama 3 Herd of Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.858777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.858777Z digest=sha256:470432126f901d47d8aace76a4097a99ee467d38fa955cbba082f2f01a6dfd86

Observation e50db835-66ed-4bc0-be9d-8dd885ffb078 · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.296489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:58.905319Z digest=sha256:ef9ad0126ba0c9bd0cf7451cc71a0b7ee7917aeaad981d367dcc9c99ca386a4d

Observation 75835048-2848-449e-b468-48470fc8d4d8 · outbound

This paper cites HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One HLM-Cite: Hybrid Language Model Workflow for Text-based Scientific Citation Prediction

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:58.977041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:58.977041Z digest=sha256:8e1fd5d25232ebd555cd733a17548bd71526bd79f70bb49e1580c191d287a2f4

Observation 208f9e6d-4ddf-402e-904d-6970c1895359 · outbound

This paper cites Stance detection with collaborative role-infused llm-based agents.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Stance detection with collaborative role-infused llm-based agents

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.274908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:59.013226Z digest=sha256:3edc278aba3381df4242a28c634ede5864022723b87a66a365ad78224a2b733c

Observation 9c59abd1-1c3a-417f-917f-fa74e16f8d15 · outbound

This paper cites A Survey of Large Language Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Survey of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.120276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.120276Z digest=sha256:5d7de41660845339a2dee1eba32b12afff0c2f69c85a0bd3ba678667241ee923

Observation 2b89717b-1b06-4ca5-a502-4431b39f1141 · outbound

This paper cites A survey on evaluation of large language models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A survey on evaluation of large language models.ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.187554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.187554Z digest=sha256:9c83af1fe07ac0a6fd7e36b529a7e64b846f9167e854f542694fe8bd48a61e31

Observation 5e1fc835-6489-454b-9871-983b131ae05f · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.239623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.239623Z digest=sha256:f5672e24f3678969b2df527c42fcc90c7fb57f16dcc83dc75cc040fc6b11e9c1

Observation 0feaeef7-b1aa-4336-a23b-c0835c105e9f · outbound

This paper cites Human-level control through deep reinforcement learning.Nature, 518(7540):529–533, 2015.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Human-level control through deep reinforcement learning.Nature, 518(7540):529–533, 2015

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.336489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.336489Z digest=sha256:5edf0cd72010fe748f194e4d52e90c0c01d4ce7046bef27c128b27a0b49b2297

Observation 64033e28-5714-48c8-ba06-23aa04f9f757 · outbound

This paper cites Deep reinforcement learning based ensemble model for rumor tracking.Information Systems, 103:101772, 2022.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Deep reinforcement learning based ensemble model for rumor tracking.Information Systems, 103:101772, 2022

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.226796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:59.406512Z digest=sha256:ad43f421dee57629f5b51fcf5c71a7b8861042700a21d9b623aa023edb5a40f7

Observation af3c0484-5644-417c-9949-b2722bfb93fe · outbound

This paper cites An oppositional-cauchy based gsk evolutionary algorithm with a novel deep ensemble reinforcement learning strategy for covid-19 diagnosis.Applied Soft Computing, 111:107675, 2021.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One An oppositional-cauchy based gsk evolutionary algorithm with a novel deep ensemble reinforcement learning strategy for covid-19 diagnosis.Applied Soft Computing, 111:107675, 2021

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.179323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:59.516768Z digest=sha256:c03ab2562239baf30bacc164a1b8a0624855d8d5a25cf2d63140d31593e9ce66

Observation fe683cb7-9a2a-4882-8a66-49e4a15e6b1d · outbound

This paper cites Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.577571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.577571Z digest=sha256:c0dedd1df4f2fd808368d59b1c72f14773102d04c254f3f18992152f3d074246

Observation d068ec1c-548d-4d1b-8f5e-9195ddd2e1f5 · outbound

This paper cites Augmenting autotelic agents with large language models.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Augmenting autotelic agents with large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:01.080388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:59.675486Z digest=sha256:11e58392d179505fc5b8a29e3077ffd48fc67ead1be7eee9a0e263b1a47afcb7

Observation 5674d13e-b4ba-43fb-b7cc-48b26df7d801 · outbound

This paper cites Read and reap the rewards: Learning to play atari with the help of instruction manuals.Advances in Neural Information Processing Systems, 36, 2024.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Read and reap the rewards: Learning to play atari with the help of instruction manuals.Advances in Neural Information Processing Systems, 36, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.823614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:59.806577Z digest=sha256:ef3fce844ad2abeb0e6ba9a650252b34409993cc71f3a2806d2d26f5a4cf0c6f

Observation 1bfd96f7-721e-46bc-821e-f6528fd842e8 · outbound

This paper cites Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Self-Refined Large Language Model as Automated Reward Function Designer for Deep Reinforcement Learning in Robotics

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:59.961027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:59.961027Z digest=sha256:d4e9bafa472e0ea1e5bfe45e5d85eb9ed0b34c6f16b2efb1a8d7343b7867c7b9

Observation e6e1b00b-aa8c-4a42-a3a7-d870bb862034 · outbound

This paper cites Text2reward: Reward shaping with language models for reinforcement learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Text2reward: Reward shaping with language models for reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.632744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:25:00.040054Z digest=sha256:503a60e7b997c513e7b228901b6b23da647c8b0a2184f2ce5385801c22f58c7f

Observation 61197aeb-f728-4ac8-bfb1-da8bba5d30b2 · outbound

This paper cites LLM-Empowered State Representation for Reinforcement Learning.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One LLM-Empowered State Representation for Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:25:00.044677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:25:00.044677Z digest=sha256:3b2f8aa9ec953cd54c5508c87f5ebb72fc09d5bcb3cde18ddb9312da15609f0b

Observation dcc67497-61f2-4cda-bb79-433cc4f087b4 · outbound

This paper cites Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:25:00.049554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:25:00.049554Z digest=sha256:1341b0be45bb0510a7e536c673e95ec389d37284668306661c866ed08c4ebfa3

Observation b0b4ac66-6707-458b-a5e8-9cec8977cf8d · outbound

This paper cites Exploration Situation.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One Exploration Situation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:25:00.520873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:25:00.054246Z digest=sha256:f2735b8669b481d2f22f2e8b0ea865450e84c711ee11036a99b6f7b70f061578

Pith citing papers

No inbound Pith citation observations are available.