Pith. sign in

Paper Citation Record · LEDGER

Great Models Think Alike and this Undermines AI Oversight

As of 9 August 2026, this Paper Citation Record lists 100 of 107 outbound references and 11 inbound Pith citation observations for arXiv:2502.04313.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04313 v2

Coverage vector

measured 100 of 107 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:54:02.909618Z

measured 111 of 111 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:34:26.874878Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:57:41.567131Z

Reference resolution

100 of 107 outbound references displayed

  • verified exact0
  • verified fuzzy61
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 351a9710-201a-4bdf-90ae-22f274d07cd7 · outbound

This paper cites write newline.

Great Models Think Alike and this Undermines AI Oversight write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.499159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.499159Z digest=sha256:b4f35df8836c09e663c44830155c8ab51583033060c8572293e3668ea6b3acec

Observation bc38d202-aded-49e1-9ff5-cd10b94b19ab · outbound

This paper cites B., Lozhkov, A., Bakouch, E., Blázquez, G.

Great Models Think Alike and this Undermines AI Oversight B., Lozhkov, A., Bakouch, E., Blázquez, G

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.505415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.505415Z digest=sha256:eed03cdf546fe50854bb6d424e4b147d80b5087568fadcda6cd9314335419869

Observation 10a20e74-d1d6-4b55-b33b-7565d8dbb01a · outbound

This paper cites and Perez-Villadoniga, M.

Great Models Think Alike and this Undermines AI Oversight and Perez-Villadoniga, M

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.510124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.510124Z digest=sha256:d03edbf005bd0ff01da3b41c69b110a71d43540e1dc55f6c01c5b363c0e84976

Observation 89062b19-45b1-4226-a52e-0cba059ca941 · outbound

This paper cites Towards evaluations-based safety cases for ai scheming, 2024.

Great Models Think Alike and this Undermines AI Oversight Towards evaluations-based safety cases for ai scheming, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.514664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.514664Z digest=sha256:61a037683bcf85322b9ba41b0f24676ef7a9736744dd277a98cc0b5e8ab12430

Observation fe68c7e8-f373-421d-840d-4987804141cd · outbound

This paper cites Revisiting model stitching to compare neural representations.

Great Models Think Alike and this Undermines AI Oversight Revisiting model stitching to compare neural representations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.518935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.518935Z digest=sha256:99d7b7b201666c3a5298a246d9d8bbea172aa8960496560600aef5c131ece852

Observation 7344e4e3-9b00-4626-90e6-7c2123e8eda2 · outbound

This paper cites F., Ammanamanchi, P.

Great Models Think Alike and this Undermines AI Oversight F., Ammanamanchi, P

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.523516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.523516Z digest=sha256:fdd9649ca8cd2a962202705b2db0b079abe0f70c5d41bae006ae2269dcd555ab

Observation 624b9ca3-c429-4534-88a9-828f4b9d4514 · outbound

This paper cites Holistic evaluation of language models.

Great Models Think Alike and this Undermines AI Oversight Holistic evaluation of language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.527965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.527965Z digest=sha256:dd366da97f2a650e77b86efa54b8f0988717dd805385d7613a2e96a38a5bb6f5

Observation d6067d91-2756-42a3-87e7-fa7e08141be2 · outbound

This paper cites Which prompts make the difference? data prioritization for efficient human llm evaluation, 2023.

Great Models Think Alike and this Undermines AI Oversight Which prompts make the difference? data prioritization for efficient human llm evaluation, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.532560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.532560Z digest=sha256:34478fd54093149c8731304f248834cea910797cfcad17ac39b3e5f5c77067c9

Observation 7ca69eac-b5cd-41c4-893b-f0049538c184 · outbound

This paper cites an unresolved cited work.

Great Models Think Alike and this Undermines AI Oversight Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.536524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.536524Z digest=sha256:83f73af404f2b5e55d3f7071d3e610eb5abd2f52e558f68928d131cfe8de4b82

Observation f29e235b-10ba-4be1-9c6f-41f213132658 · outbound

This paper cites D., Martinez-Plumed, F., Tenenbaum, J.

Great Models Think Alike and this Undermines AI Oversight D., Martinez-Plumed, F., Tenenbaum, J

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.540744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.540744Z digest=sha256:0f16137e67b2f6acf3a986ee4252892b4fd1b8c18f60182a6679fcedb74de631

Observation e5b1e0a9-87f3-4ad6-8b5e-275672afead2 · outbound

This paper cites H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., Sutskever, I., and Wu, J.

Great Models Think Alike and this Undermines AI Oversight H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., Sutskever, I., and Wu, J

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.544921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.544921Z digest=sha256:36c887760dde8feacbdf9fce911d92dd4137b0e5bc635d0819a0f4de5ba99fca

Observation 5463a38a-151a-42ab-911c-99f1b01f1435 · outbound

This paper cites A portfolio approach to research funding.

Great Models Think Alike and this Undermines AI Oversight A portfolio approach to research funding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.549028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.549028Z digest=sha256:627de7ffa6c6b3449a9c0279eaba904dcf7c4293311670446e54e274a9396da5

Observation 32d9e6ff-23c4-407f-8081-3afa6e918316 · outbound

This paper cites Quantifying the gain in weak-to-strong generalization.

Great Models Think Alike and this Undermines AI Oversight Quantifying the gain in weak-to-strong generalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.553066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.553066Z digest=sha256:3b3a5f6ab3aed13f3de0beb3b33b2fc2938d204e3d7fe6e3272704ee9a63cd74

Observation 084cabde-0eff-45f9-aeaa-1edf959799c1 · outbound

This paper cites H., Chen, S., Liu, Z., Jiang, F., and Wang, B.

Great Models Think Alike and this Undermines AI Oversight H., Chen, S., Liu, Z., Jiang, F., and Wang, B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.557083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.557083Z digest=sha256:60c83ac93a407ab3dcc2c8f1a30ef6297429bec95454e57b4099786d53913d81

Observation 2a37be34-8b04-4434-b4e8-288471cf7a27 · outbound

This paper cites E., Stoica, I., and Xing, E.

Great Models Think Alike and this Undermines AI Oversight E., Stoica, I., and Xing, E

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.561178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.561178Z digest=sha256:7e0761edefe77f2b748d086874b3df6fd85a3c8ac1dcc77bf7ef9e41e0c40bfd

Observation 24d6c2de-4546-48b9-aef6-256df238b6e4 · outbound

This paper cites J., and Jurman, G.

Great Models Think Alike and this Undermines AI Oversight J., and Jurman, G

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.565311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.565311Z digest=sha256:d9766c85c633872e5475a6ae05131cfa93402f97bfd9de828f6126d7879b5e6e

Observation ecd41ed7-e481-4083-9296-187d7425cec7 · outbound

This paper cites B ool Q : Exploring the surprising difficulty of natural yes/no questions.

Great Models Think Alike and this Undermines AI Oversight B ool Q : Exploring the surprising difficulty of natural yes/no questions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.569270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.569270Z digest=sha256:ea20ddff9092c98b0005295a921585a23ce90cf2c9d7c77bcfb5da175635e776

Observation 28e6491f-b70e-4aad-a32a-8e60fa5b7d8b · outbound

This paper cites A coefficient of agreement for nominal scales.

Great Models Think Alike and this Undermines AI Oversight A coefficient of agreement for nominal scales

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.573510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.573510Z digest=sha256:56e7d50baad5153c882629ede01759daf8cd9ca717a7ea1a9b3fc1f9e6b349da

Observation d5b12e98-b4f1-4e91-b267-980bbb025e79 · outbound

This paper cites X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P.

Great Models Think Alike and this Undermines AI Oversight X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.577701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.577701Z digest=sha256:e373a4fcad459edc19732f74016f223d2a848cbe9e4eb2ca4991a98fe09f1eab

Observation 1101ec06-8c75-45ee-bf4a-28395cdcc416 · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators.

Great Models Think Alike and this Undermines AI Oversight Length-controlled alpacaeval: A simple debiasing of automatic evaluators

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.581768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.581768Z digest=sha256:379632cb06700f0379aa5bfc409f646b93a60c27a2e73801269f97520c6d72b2

Observation 75528b9c-8fd8-474f-9c6e-ae9f344c4aaf · outbound

This paper cites E., and Yeung-Levy, S.

Great Models Think Alike and this Undermines AI Oversight E., and Yeung-Levy, S

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.585709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.585709Z digest=sha256:b56ce089d087f337f54f8d3f977d8b3f9504d8349ca9408084964199381961e6

Observation a66b2253-dc09-4891-91e9-1ced40e36761 · outbound

This paper cites an unresolved cited work.

Great Models Think Alike and this Undermines AI Oversight Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.589693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.589693Z digest=sha256:ddcd922283da1d3b55faa0603931cbdda0a6bac76a093c3bfb7937e6889ee6ba

Observation 4afd2635-189d-4323-b26e-97f240d96513 · outbound

This paper cites Accuracy is not all you need.

Great Models Think Alike and this Undermines AI Oversight Accuracy is not all you need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.593501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.593501Z digest=sha256:25fc0182c452e1ae1ee0f39dfd10029fe6e26b4f4c6aa5676f83994076e7f1cd

Observation a35b7e18-5bfd-4714-ba92-8451c889f33e · outbound

This paper cites Model changelists: Characterizing updates to ml models.

Great Models Think Alike and this Undermines AI Oversight Model changelists: Characterizing updates to ml models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.597561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.597561Z digest=sha256:90c7546a3565f997036ddb68bbfa02e2bd88966f3f91c323eea595878d8c4a3e

Observation 5d3ca6ba-958d-43fd-b3fe-1a82d3238d80 · outbound

This paper cites L., Levin, B., Paik, M.

Great Models Think Alike and this Undermines AI Oversight L., Levin, B., Paik, M

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.601484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.601484Z digest=sha256:dcf981d6beab23bda169198caa9f5ffb2d25e4794c4c77d3984e3e4725e502f9

Observation fdf8696a-8292-459a-b836-9c3b6403c427 · outbound

This paper cites Evaluating superhuman models with consistency checks.

Great Models Think Alike and this Undermines AI Oversight Evaluating superhuman models with consistency checks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.605595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.605595Z digest=sha256:f1873d07abab2dc35ab4dd8aa194e11fd01d0f9706f56427baff29afe6acd1a6

Observation 88e5db68-1477-4769-a059-f159ed67112d · outbound

This paper cites A framework for few-shot language model evaluation, 2023.

Great Models Think Alike and this Undermines AI Oversight A framework for few-shot language model evaluation, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.609629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.609629Z digest=sha256:270bd5ce875f8cae6cfa0417077ad379f879ea0acb556e2fa1a639dc2d7d2440

Observation 22b7d1ca-91a2-49cc-80c8-1677fbf80df5 · outbound

This paper cites an unresolved cited work.

Great Models Think Alike and this Undermines AI Oversight Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.613633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.613633Z digest=sha256:5f634b9528b27e327b1b10f5d2ce2750afed177b24ceec09e870502e7a267db3

Observation 56bb0f04-c84f-4d33-9564-4ebbf35f5235 · outbound

This paper cites A., and Brendel, W.

Great Models Think Alike and this Undermines AI Oversight A., and Brendel, W

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.218211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.617652Z digest=sha256:37d64910caf892a365fa16d19f653bc969e63da7a2f49a5fc21ec253d6a6bae2

Observation 2fa2f3f2-dc7a-49eb-849f-2439bdb693aa · outbound

This paper cites Gemma 2: Improving open language models at a practical size, 2024.

Great Models Think Alike and this Undermines AI Oversight Gemma 2: Improving open language models at a practical size, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.204866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.621699Z digest=sha256:a6dea0ce1062e140050251ef6304fb2b3df12025ce2e38de8dbbe0c7aa50fae2

Observation 91b98932-cf8c-4aad-89b9-5ba30303ba90 · outbound

This paper cites Onebench to test them all: Sample-level benchmarking over open-ended capabilities, 2024.

Great Models Think Alike and this Undermines AI Oversight Onebench to test them all: Sample-level benchmarking over open-ended capabilities, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.191341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.625821Z digest=sha256:a4564dcfe1f2118fb115d8c6ce9cd375b8cde1f89d0b7da1d5ef03788821dd97

Observation 1f06f59e-2779-469f-a538-15c138c5c6d6 · outbound

This paper cites Chatgpt outperforms crowd workers for text-annotation tasks.

Great Models Think Alike and this Undermines AI Oversight Chatgpt outperforms crowd workers for text-annotation tasks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.177199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.629816Z digest=sha256:2f8f967eb19ba35cd4c083e3431920478f7255dce4c8fde336d94239f317f3a5

Observation fc14ef14-2eaa-4b61-a0ab-a8ed12c72de8 · outbound

This paper cites and Dao, J.

Great Models Think Alike and this Undermines AI Oversight and Dao, J

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.163478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.633806Z digest=sha256:0eb759db70289bfa7cc71be99a7d4c42d09ef5e9313e9401bf5e1ec3af0b38e0

Observation a427666a-3b47-4300-b504-162d2bf457e2 · outbound

This paper cites and Dao, T.

Great Models Think Alike and this Undermines AI Oversight and Dao, T

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.637905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.637905Z digest=sha256:6a0f87f4523aac854c1b8d51fd97e454585430b006b30e71e54370ace19a439f

Observation 1523b8ad-7dde-4f9c-95a1-260b23b97278 · outbound

This paper cites Vision superalignment: Weak-to-strong generalization for vision foundation models, 2024.

Great Models Think Alike and this Undermines AI Oversight Vision superalignment: Weak-to-strong generalization for vision foundation models, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.140966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.641949Z digest=sha256:c328faf9b8c508b267006a66d16cf00ef01ceff4c819633e40fc1d5cf9719cc7

Observation 577b46d6-cd14-491c-939b-27ca786e87a2 · outbound

This paper cites On the blind spots of model-based evaluation metrics for text generation.

Great Models Think Alike and this Undermines AI Oversight On the blind spots of model-based evaluation metrics for text generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.127725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.645865Z digest=sha256:c3202474161e373ad758f4c394217d83bfb1aa5868570aa750507623971277dc

Observation aaca59ab-f8ae-4b2e-ac1e-4c3e91210d95 · outbound

This paper cites Aligning AI with shared human values.

Great Models Think Alike and this Undermines AI Oversight Aligning AI with shared human values

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.114371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.649760Z digest=sha256:af0aff22bf95a816f160392f39cddcd564150dcf2424ab5411f85d092a95e71b

Observation 011a5697-ca30-46e0-8686-78f9ab291e95 · outbound

This paper cites Measuring massive multitask language understanding.

Great Models Think Alike and this Undermines AI Oversight Measuring massive multitask language understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.100840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.653975Z digest=sha256:88bc73598ff98aba4d6a58a4070ae8f64bd1d3888c4099635a63d13b03ead8ff

Observation e64800b2-7fa1-459f-aaae-e297f456c768 · outbound

This paper cites J., Shen, Y., Wallis, P., Allen - Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

Great Models Think Alike and this Undermines AI Oversight J., Shen, Y., Wallis, P., Allen - Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.086162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.657997Z digest=sha256:5f59719ffad3317a39a0b593b7abed48f2780be1c84be2be8ad0abc3af3546f0

Observation 9ddc60a5-eb3b-4892-8166-edfe57ce0369 · outbound

This paper cites Cosmos QA : Machine reading comprehension with contextual commonsense reasoning.

Great Models Think Alike and this Undermines AI Oversight Cosmos QA : Machine reading comprehension with contextual commonsense reasoning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.072418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.662004Z digest=sha256:bb1302193b26ab2fd761060ed8892e7ebfafd08c339bce0ca95f4633d14629a0

Observation f6919b46-2428-4dab-aadb-e1177b86fac4 · outbound

This paper cites D., Parker-Holder, J., Behbahani, F., Mavalankar, A., Shi, Y., Schaul, T., and Rockt\" a schel, T.

Great Models Think Alike and this Undermines AI Oversight D., Parker-Holder, J., Behbahani, F., Mavalankar, A., Shi, Y., Schaul, T., and Rockt\" a schel, T

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.058759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.666026Z digest=sha256:ac8a344513586b6748cfb7541a1c5b5396fb4b31e5af409fcb66f42465ff16bb

Observation 14bca023-6aee-4e68-9801-66c0a6bcfb53 · outbound

This paper cites Position: the platonic representation hypothesis.

Great Models Think Alike and this Undermines AI Oversight Position: the platonic representation hypothesis

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.045168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.670428Z digest=sha256:f58b6a8d883f79cc160390c7097d90ea45e4cdaf5cd3c5b83fc7398a82b058ca

Observation 5396c51a-4ce2-4242-858e-762990d8d36c · outbound

This paper cites X., Wexler, J., Reif, E., Kallarackal, K., Chang, M., Terry, M., and Dixon, L.

Great Models Think Alike and this Undermines AI Oversight X., Wexler, J., Reif, E., Kallarackal, K., Chang, M., Terry, M., and Dixon, L

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.031544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.674409Z digest=sha256:32ea87b9f2cdddc52051e4a3e2983e199fa986fc3fa6a873b9d2c1f8c96371bc

Observation 2aca9ea4-44af-4917-afe9-d1cf0782afd5 · outbound

This paper cites B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D.

Great Models Think Alike and this Undermines AI Oversight B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.678656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.678656Z digest=sha256:006057bfdd8bf9087bcf74c76fc57bffc00e07a0430cb56d4ee823e25820deb4

Observation acc281c0-7206-4cb5-8003-b24bf1fa8bf6 · outbound

This paper cites L., and Koyejo, S.

Great Models Think Alike and this Undermines AI Oversight L., and Koyejo, S

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:04.009467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.682664Z digest=sha256:bf690f608df2d9811eeb05f5b7e8a4eed787221dcc0c0e7df9812437f5871f5c

Observation bc67f1dd-d232-4ef3-b008-d546ed07c630 · outbound

This paper cites Y., Kram\' a r, J., Brown-Cohen, J., Albanie, S., Bulian, J., Agarwal, R., Lindner, D., Tang, Y., Goodman, N.

Great Models Think Alike and this Undermines AI Oversight Y., Kram\' a r, J., Brown-Cohen, J., Albanie, S., Bulian, J., Agarwal, R., Lindner, D., Tang, Y., Goodman, N

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.996002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.686682Z digest=sha256:64e215a59f62706e297c0a96eba0daf3cbdf7bf57d8b9a11af16c69a8031e79f

Observation 12504cf5-ffd8-44a4-b842-26f512d89830 · outbound

This paper cites Looking beyond the surface: A challenge set for reading comprehension over multiple sentences.

Great Models Think Alike and this Undermines AI Oversight Looking beyond the surface: A challenge set for reading comprehension over multiple sentences

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.982653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.690615Z digest=sha256:f40eee6f8e457d6ae1e418cac39073687bac891373c0dc7ad1a6fdf2b4bbd21b

Observation d274a51d-3517-4daf-803b-99267822a003 · outbound

This paper cites Similarity of neural network models: A survey of functional and representational measures.

Great Models Think Alike and this Undermines AI Oversight Similarity of neural network models: A survey of functional and representational measures

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.969140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.694654Z digest=sha256:32aea398720b7f846536467ec7cb69916962c9b9f37755fe09dc9362e91000e9

Observation 99200a00-7cf0-454f-bc8b-e4f7501e1275 · outbound

This paper cites and Raghavan, M.

Great Models Think Alike and this Undermines AI Oversight and Raghavan, M

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.698830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.698830Z digest=sha256:88e941b7132fc0fde6c0001cdfad5eb6201cdf341605ba170766cfd5c3a008c2

Observation cec46f14-7674-4a28-935a-2ca32d7d5ad5 · outbound

This paper cites To ship or not to ship: An extensive evaluation of automatic metrics for machine translation.

Great Models Think Alike and this Undermines AI Oversight To ship or not to ship: An extensive evaluation of automatic metrics for machine translation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.946538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.703251Z digest=sha256:1e8a8208df3ada06dd0e5fe0b8bcd976ba849f2087a254416449b915fe5ae493

Observation 05fb5b76-13a7-444a-9df8-107308c4f602 · outbound

This paper cites R., Vaidya, A., Mahmood, F., Zitnik, M., Chen, T., and Hartvigsen, T.

Great Models Think Alike and this Undermines AI Oversight R., Vaidya, A., Mahmood, F., Zitnik, M., Chen, T., and Hartvigsen, T

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.932813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.707242Z digest=sha256:55d630dedaadbf361bd8c97f0b05854417f993711e97df09c47ddeb81a03a9ed

Observation e009bac9-3949-4fc9-9c3d-f15ad50d6fa7 · outbound

This paper cites I., Kim, Z.

Great Models Think Alike and this Undermines AI Oversight I., Kim, Z

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.919152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.711299Z digest=sha256:068371fd0f17e43719d7f1aa9d9bcdc7fa03944bebca2308131d5217c8f5e3ed

Observation aae7df1f-99f2-4617-b531-964b0cf21a23 · outbound

This paper cites Similarity of neural network representations revisited.

Great Models Think Alike and this Undermines AI Oversight Similarity of neural network representations revisited

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.905440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.715539Z digest=sha256:3f4274676f365e80d17cfe79cb465bf0e05b142ce6bab732bf668f33f8a4dbed

Observation 391c8edd-cc8f-494c-8082-bc183cfc1955 · outbound

This paper cites Reliability in content analysis: Some common misconceptions and recommendations.

Great Models Think Alike and this Undermines AI Oversight Reliability in content analysis: Some common misconceptions and recommendations

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.891595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.719705Z digest=sha256:eb6d42807f0b400bd3a85aacc0f772fb58e55941734d94cae93f432f998eb800

Observation 6669ddae-4019-4e0c-9618-21c3daf19486 · outbound

This paper cites H., Gonzalez, J.

Great Models Think Alike and this Undermines AI Oversight H., Gonzalez, J

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.723706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.723706Z digest=sha256:b60db950275c208f61fceae0a213fc8ff56cc21f590afcb029a997bedcaff41f

Observation aff27eea-b36d-49a2-9410-1c9eac791ca2 · outbound

This paper cites D., Dombrowski, A.-K., Goel, S., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A.

Great Models Think Alike and this Undermines AI Oversight D., Dombrowski, A.-K., Goel, S., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.869351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.727875Z digest=sha256:54e0761c98737a1af16ede2dfbbd85e652ecf4795ad87668a9e7e9db20f36f64

Observation ef499afc-6151-4c0c-af17-1976f19a17a1 · outbound

This paper cites E., and Stoica, I.

Great Models Think Alike and this Undermines AI Oversight E., and Stoica, I

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.856488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.731876Z digest=sha256:4d67f0e300e0f4ea8c800368628bda7582e3d23abecff18de6bc05fa99ad8a82

Observation da7f1fe0-081d-4e5e-a00c-9788ef24291b · outbound

This paper cites D., Gunasekar, S., and Lee, Y.

Great Models Think Alike and this Undermines AI Oversight D., Gunasekar, S., and Lee, Y

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.843754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.736143Z digest=sha256:04ce6031224106ebe743490185fa83632224037b173340e3a71f06b16fbf044e

Observation 0087be4f-4050-4954-ab04-ef6a462875cc · outbound

This paper cites Let's verify step by step.

Great Models Think Alike and this Undermines AI Oversight Let's verify step by step

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.830938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.740279Z digest=sha256:ef0142499975b06952c219acb8b714500f5dbd8062a5295835f53d6268243b45

Observation 3fee68f5-8e25-4a7d-a80a-9fb13686e5c5 · outbound

This paper cites LLM s as narcissistic evaluators: When ego inflates evaluation scores.

Great Models Think Alike and this Undermines AI Oversight LLM s as narcissistic evaluators: When ego inflates evaluation scores

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.818229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.744368Z digest=sha256:fb38547482385d62512bbd02ce479ddd1efb8a7991805f5621ba293b17c62fca

Observation 19826aad-942b-497f-852b-3ab3ced49dff · outbound

This paper cites The llama 3 herd of models, 2024 a.

Great Models Think Alike and this Undermines AI Oversight The llama 3 herd of models, 2024 a

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.805223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.748523Z digest=sha256:2bd5ade166878fe3715e1866225668bebc6160887f34e5da8d59c6f406fa8ade

Observation f42c26ec-d16a-44df-99c2-f4f7be20b5e4 · outbound

This paper cites Llama 3.2 model card.

Great Models Think Alike and this Undermines AI Oversight Llama 3.2 model card

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.792392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.752626Z digest=sha256:44ab6d7f3c4ccaaef42dae1f2ccd166aa62b80b4db7e1baf4505c4c656011fda

Observation 821877d2-719d-4546-ae55-0a1ac1066388 · outbound

This paper cites Llama 3.3 model card.

Great Models Think Alike and this Undermines AI Oversight Llama 3.3 model card

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.779654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.756683Z digest=sha256:39e18ed5d72bda75aeffcf39b801b6f5bbe9df98c3d2ad4efbcbe3a174fef016

Observation 950325e3-7131-4f2e-ad0f-b84082a1026e · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments.

Great Models Think Alike and this Undermines AI Oversight Multi-agent actor-critic for mixed cooperative-competitive environments

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.766461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.760827Z digest=sha256:921a4f30e057513a00400fbf98499d0948c54cb18de43f1ebce6cc4cc103473f

Observation 37d5a215-6c58-4727-8221-83a378d7ea81 · outbound

This paper cites An adversarial perspective on machine unlearning for AI safety.

Great Models Think Alike and this Undermines AI Oversight An adversarial perspective on machine unlearning for AI safety

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.753002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.764898Z digest=sha256:92791983cfb03a09cd3f5147eb5c48b6133ad76a071abf8a99ba2591e9ac05c1

Observation 026067bc-61f9-4dcf-b7df-c9d80e3cc79b · outbound

This paper cites Aidanbench: Stress-testing language model creativity on open-ended questions.

Great Models Think Alike and this Undermines AI Oversight Aidanbench: Stress-testing language model creativity on open-ended questions

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.738463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.771222Z digest=sha256:714651e9ad387613a884750193f3f692005b325e8ddf3ad2d7a608ee97d5b9ad

Observation 36930aaa-6753-4d3b-805a-242a039a054d · outbound

This paper cites Phi-4 technical report.

Great Models Think Alike and this Undermines AI Oversight Phi-4 technical report

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.724854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.775437Z digest=sha256:14606e71d2384bfb28f8090af1941ac96ddf390e64bbcf930c0d6ae164004573

Observation ae7f2323-a1a7-40d0-a21c-26f49cd5ea7a · outbound

This paper cites Ministral 8b instruct model card.

Great Models Think Alike and this Undermines AI Oversight Ministral 8b instruct model card

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.711740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.779571Z digest=sha256:f869995b2bfeaa7519b186b367d897d6a48ccb754214cf98e3d81ada6b44a935

Observation 56b076bc-a7b1-429d-8f48-c2edc75bf658 · outbound

This paper cites M., and Shen, Z.

Great Models Think Alike and this Undermines AI Oversight M., and Shen, Z

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.698254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.783597Z digest=sha256:6b7a36e0245bbfec4af4bcef5bea7ee10f78762015300e86320846046a30c8a6

Observation 4f2031d4-2663-4d7e-a8de-6f0f128cad89 · outbound

This paper cites Adversarial NLI : A new benchmark for natural language understanding.

Great Models Think Alike and this Undermines AI Oversight Adversarial NLI : A new benchmark for natural language understanding

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.684699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.787661Z digest=sha256:b4cd6f8c29e12b46890a813ea2708d3573869d521842e2e0aba4bf7ec86d574b

Observation 328b9d0d-6cc0-4bbb-bd07-79496927ccad · outbound

This paper cites Gpt-4 technical report, 2024.

Great Models Think Alike and this Undermines AI Oversight Gpt-4 technical report, 2024

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.791538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.791538Z digest=sha256:e1359dfbee8a50f60e4dc19ccaf402883ca9cb13d962b1420b7ef3f8aa2ca28e

Observation 66d365a5-05ea-4b03-aaaa-4062a962255a · outbound

This paper cites F., Leike, J., and Lowe, R.

Great Models Think Alike and this Undermines AI Oversight F., Leike, J., and Lowe, R

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.661452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.795542Z digest=sha256:587561ef851af247cd39579e64f786e4b79401f208d3bb8b838a0f51ab247665

Observation 3a369e5c-8f59-4496-8690-319492825dae · outbound

This paper cites R., and Feng, S.

Great Models Think Alike and this Undermines AI Oversight R., and Feng, S

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.646836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.799619Z digest=sha256:03a7ac092054d29acdc22f6340ac3db76ec46f9717b16b6de20bf4413421ebca

Observation 2abf2ede-959b-422a-88ff-419eda4c13b8 · outbound

This paper cites B leu: a method for automatic evaluation of machine translation.

Great Models Think Alike and this Undermines AI Oversight B leu: a method for automatic evaluation of machine translation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.633349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.803793Z digest=sha256:80a2b8575313386dbe1933000286b918ec0ab88d42c29c8a802b2896d50b51ab

Observation ee4d54a4-1ea5-4858-b3d6-eea30d8a629e · outbound

This paper cites an unresolved cited work.

Great Models Think Alike and this Undermines AI Oversight Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:54:03.619371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.807849Z digest=sha256:26922fc243daae52aeb1e71c186f6842654b6588ab0bd78863cace13cce0fbc7

Observation ad3824cc-3727-4ee6-89ab-6c116585eff8 · outbound

This paper cites Mauve: Measuring the gap between neural text and human text using divergence frontiers.

Great Models Think Alike and this Undermines AI Oversight Mauve: Measuring the gap between neural text and human text using divergence frontiers

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.605737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.811963Z digest=sha256:75649ddeff6bb2e7924c5ebd7ab544a773356973101e29a916bf01f57b2c1fb2

Observation 9309217f-4d70-46d3-95ee-d8e5eb3237d7 · outbound

This paper cites Qwen2.5 technical report, 2025.

Great Models Think Alike and this Undermines AI Oversight Qwen2.5 technical report, 2025

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.591495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.815926Z digest=sha256:99a2dc0d601543f13c1a962ba9b599da2d22fcc2d27c3939fb51bedc64a52fd7

Observation cab7d1eb-4878-4b71-8bd1-b0bffb28679e · outbound

This paper cites Language models are unsupervised multitask learners, 2019.

Great Models Think Alike and this Undermines AI Oversight Language models are unsupervised multitask learners, 2019

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:02.819965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:54:02.819965Z digest=sha256:d1d71c88d326e68a5a19ef37d118c45f56585bab06db737c3e06273107ceeb46

Observation 1e90a77f-c52c-4d8d-8f21-eb322a4ed88e · outbound

This paper cites Getting closer to ai complete question answering: A set of prerequisite real tasks.

Great Models Think Alike and this Undermines AI Oversight Getting closer to ai complete question answering: A set of prerequisite real tasks

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.569218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.823883Z digest=sha256:602281855b14e86a39b4d857ef345e157ad17741620eaaed8edefdc6a87c2fdd

Observation afb1a7e9-d819-4631-975c-b82f00feb3ec · outbound

This paper cites S., Vinyals, O., H \' e naff, O.

Great Models Think Alike and this Undermines AI Oversight S., Vinyals, O., H \' e naff, O

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.555784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.827877Z digest=sha256:8b3e1b8b4a00639d8ab68361c13b8b1ebd8d926b5ad51ff58748e2813493bf67

Observation d5265bf4-8d01-498f-9ec9-2e1da7b37e89 · outbound

This paper cites Min-mid-max scaling, limits of agreement, and agreement score, 2020.

Great Models Think Alike and this Undermines AI Oversight Min-mid-max scaling, limits of agreement, and agreement score, 2020

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.542611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.831879Z digest=sha256:583c93b75fedf4812752f1080c97928da742b7308e5859895a5e2721fc5118a2

Observation 82105bf6-11b5-4f56-9a45-52dcc379f91f · outbound

This paper cites Social IQ a: Commonsense reasoning about social interactions.

Great Models Think Alike and this Undermines AI Oversight Social IQ a: Commonsense reasoning about social interactions

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.529048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.836051Z digest=sha256:05be0f60ded9dbd135770f14023998e810c8344ad8bd4addc07b763defc65134

Observation 938acea5-090e-442f-9df2-97cfa69d2e05 · outbound

This paper cites Experiments in weak-to-strong generalization, 2024.

Great Models Think Alike and this Undermines AI Oversight Experiments in weak-to-strong generalization, 2024

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.514892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.840396Z digest=sha256:6bad47e86dca1c83f18d8015af0fb232735a14c5fc0ae86493fe680dbffa1807

Observation efea378a-720d-4c2e-bf35-01fe8e8ba077 · outbound

This paper cites an unresolved cited work.

Great Models Think Alike and this Undermines AI Oversight Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:54:03.501919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.844416Z digest=sha256:060c5de76cf124f3f3893192321858c3078eef8a6bfd2816049586cd7de13dea

Observation f3f392a7-9ab5-445c-8ec0-f4d918b80868 · outbound

This paper cites M., Ilyas, A., and Madry, A.

Great Models Think Alike and this Undermines AI Oversight M., Ilyas, A., and Madry, A

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.488787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.848711Z digest=sha256:11d9269f696ff6248f25799e03f8b59bf8d1353a46095764d02a384cebcbacc5

Observation f8b9f661-c9e6-4a4c-a1ce-c68cda52ccf3 · outbound

This paper cites an unresolved cited work.

Great Models Think Alike and this Undermines AI Oversight Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:54:03.475762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.852621Z digest=sha256:538fcf1d7662f7cf14c81002b6950a8386a0bbc94c08c229150468950bd0bf46

Observation b0e43dfe-b66a-4665-9c56-748ef334c8f5 · outbound

This paper cites D., Ng, A., and Potts, C.

Great Models Think Alike and this Undermines AI Oversight D., Ng, A., and Potts, C

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.462581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.856798Z digest=sha256:e87c1a7a50ddae6c7908fd77bc8e180f56cf9a897d35113c08a020e991bd17c2

Observation 5df0d541-d1e1-49b5-a558-dc3bc8901fbf · outbound

This paper cites M., Foster, D.

Great Models Think Alike and this Undermines AI Oversight M., Foster, D

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.449366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.860835Z digest=sha256:01e61d99b8e0017e7956d69a3e270aecc8d57d063e8a2e8b78a50396c530b096

Observation 8b092693-06ff-493b-a23c-d95c9e17e14c · outbound

This paper cites an unresolved cited work.

Great Models Think Alike and this Undermines AI Oversight Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:54:03.436564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.864912Z digest=sha256:79ba8d434685287d91eac939f13f49582c45b63d04236896174f0c85656d88a6

Observation 70e8f388-d5ae-4f1f-846b-18467aa59dc3 · outbound

This paper cites LM diff: A visual diff tool to compare language models.

Great Models Think Alike and this Undermines AI Oversight LM diff: A visual diff tool to compare language models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.423769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.869195Z digest=sha256:093f76a98b7be7756af17b9d7cf7b7cd2abdc26ee73837e44b7fece8f73b6396

Observation f70b4d23-7cac-4443-bf7d-733986b7df0e · outbound

This paper cites DREAM : A challenge data set and models for dialogue-based reading comprehension.

Great Models Think Alike and this Undermines AI Oversight DREAM : A challenge data set and models for dialogue-based reading comprehension

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.410743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.873174Z digest=sha256:5d581dd2ca3bc6d5a9cba882cb7c0e6517c3b82df263c184d4b5bb9317079fdc

Observation 283de731-c857-40f4-ad07-0305d81f51cd · outbound

This paper cites W., Chowdhery, A., Le, Q., Chi, E., Zhou, D., and Wei, J.

Great Models Think Alike and this Undermines AI Oversight W., Chowdhery, A., Le, Q., Chi, E., Zhou, D., and Wei, J

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.397519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.877104Z digest=sha256:bf13b4a5f7385669af545697295010bf004f1016edf50c72cf71d64ad6b5a4c4

Observation d05074b9-864b-40cd-b869-560bc18a9a4c · outbound

This paper cites Q ua RT z: An open-domain dataset of qualitative relationship questions.

Great Models Think Alike and this Undermines AI Oversight Q ua RT z: An open-domain dataset of qualitative relationship questions

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.383859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.881104Z digest=sha256:a7d1a0320118468937b10f53b1fdce59d04afe8f42fef6ac226d554774ea7c1b

Observation dd1c887c-a042-4d1b-a9d5-c857dec06772 · outbound

This paper cites Welcome to the falcon 3 family of open models! https://huggingface.co/blog/falcon3, 2024.

Great Models Think Alike and this Undermines AI Oversight Welcome to the falcon 3 family of open models! https://huggingface.co/blog/falcon3, 2024

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.370739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.885206Z digest=sha256:427e2078b7a39abb54c709531b1bab566ece199ade572beb59095a58538fdd19

Observation b2c0fee5-907d-48e7-b130-07ac9efe627d · outbound

This paper cites S., Choudhary, K., Ramayapally, V.

Great Models Think Alike and this Undermines AI Oversight S., Choudhary, K., Ramayapally, V

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.357673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.889301Z digest=sha256:43d64a9241e0b7fce1110b4027387a2f9ba4baf342fbbfec45889086720a61ab

Observation 4bf416c1-7375-4c57-8273-ade4920a50d8 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Great Models Think Alike and this Undermines AI Oversight Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.344196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.893268Z digest=sha256:1da7d0abac412c20e44ad1b79ffb816f1f825746cb13b83dc715bc003f7a48fa

Observation d79c13a6-ca4b-43ee-8413-c0b9691980bd · outbound

This paper cites an unresolved cited work.

Great Models Think Alike and this Undermines AI Oversight Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:54:03.330917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.897375Z digest=sha256:e8d7d8d0a6df527a57d3c422ae7c1b7f5afaec459ceeefd2d7f66781d8ff81d1

Observation 463ed8af-0820-4c95-b671-c07bcdb75b9a · outbound

This paper cites F., and Gardner, M.

Great Models Think Alike and this Undermines AI Oversight F., and Gardner, M

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.317385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.901522Z digest=sha256:5ac932a7099d0e03243d725fc333104ca263413c85a10a6ab1360a3f3b70adfd

Observation 25f900f7-e446-425a-82b0-6b5fe266991b · outbound

This paper cites V., and Zhang, X.

Great Models Think Alike and this Undermines AI Oversight V., and Zhang, X

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.304146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.905477Z digest=sha256:8f85fa89d66ef3140b39e0f87e680cc3f7c19109ca9f3f43dcc61d5aec1f200f

Observation def79eb0-a4ff-49cf-91a4-560bce1c66fc · outbound

This paper cites L., Tambe, M., Kakade, S., and Malach, E.

Great Models Think Alike and this Undermines AI Oversight L., Tambe, M., Kakade, S., and Malach, E

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T22:54:03.291186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:54:02.909618Z digest=sha256:54b2a3f803ac7d4ed623ec6073f7440a93018b2290fb0051b9b6e137a6cdbef2

Pith citing papers

Observation 2f377797-d54a-44d6-9f25-2988ae19ae4c · inbound

Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic Dimension cites this paper.

Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic Dimension Great Models Think Alike and this Undermines AI Oversight

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:25:20.625986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T03:24:07.851782Z digest=sha256:f06909c5892c6cdc1a11b191943008dc18e173365a1ccf7be0cda8056779959c

Observation f75e2726-15af-4092-bb23-e4c5e4ba78b6 · inbound

How Benchmark Prediction from Fewer Data Misses the Mark cites this paper.

How Benchmark Prediction from Fewer Data Misses the Mark Great Models Think Alike and this Undermines AI Oversight

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.874878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.874878Z digest=sha256:c292aec7582ff71f37ba0055915413288b829244d056f9acd358d0046be053ad

Observation 7035dbe6-4d82-4cdc-adf5-bbb23b52220d · inbound

Correlated Errors in Large Language Models cites this paper.

Correlated Errors in Large Language Models Great Models Think Alike and this Undermines AI Oversight

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.244885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.244885Z digest=sha256:bc86a0193e62610a4b95e8168bd8bf725d4166562b71ece01188a70db4b0b626

Observation 60a4a7cd-7bd7-49e7-a7e4-ff58b5ab653d · inbound

Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models cites this paper.

Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models Great Models Think Alike and this Undermines AI Oversight

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:50:59.795612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:47:31.492842Z digest=sha256:8eb0394a19cb5b67caaac048a9e1d471a9093f019f8737a6cdb6ae9eaf9ec3b9

Observation 457773f8-fc99-4712-9967-25e6931224a1 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Great Models Think Alike and this Undermines AI Oversight

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.568634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:934af93e5d51f48c7850a535f940b8034683497e651fda6b43c9f86b42fba432

Observation 521028b7-62a6-42ed-a508-2e6205bcf38b · inbound

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling cites this paper.

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling Great Models Think Alike and this Undermines AI Oversight

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:44:37.120104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T09:44:27.786630Z digest=sha256:e447996695bd78d12a1137288377a9dee37637fbd833e519eec285e744517c3b

Observation 5efa024b-040e-406c-b4e5-f7d0020ca218 · inbound

Weak-to-Strong Learning in Decision Making cites this paper.

Weak-to-Strong Learning in Decision Making Great Models Think Alike and this Undermines AI Oversight

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-01T15:26:11.571396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:26:11.571396Z digest=sha256:5ce4d46090358182343599605e2637c1b03c7aade037fbd77bc5241398173b45

Observation 73078668-45a8-44b3-aa29-74ed547193f8 · inbound

Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles cites this paper.

Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles Great Models Think Alike and this Undermines AI Oversight

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T09:29:49.269954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:29:49.269954Z digest=sha256:e4aa1f6e1214673bf5219e58b7dd6880d5566c5fa8708a90cff321feba6fa0f2

Observation 5edd23c8-e32f-443d-abcc-7c43b8f7eec3 · inbound

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop cites this paper.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Great Models Think Alike and this Undermines AI Oversight

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.784183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.784183Z digest=sha256:090ce3610d5b48d5323823c699f60e5c6a1298088cb01833a08586a6a49a283e

Observation 9664da98-6e47-449e-ac37-88e30918f897 · inbound

Language Models Agree With Each Other, Not With Readers cites this paper.

Language Models Agree With Each Other, Not With Readers Great Models Think Alike and this Undermines AI Oversight

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T10:20:13.027463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:20:13.027463Z digest=sha256:c5e282d01169b03e148fd95ea8f162d47e4ff308f4d8652221ac2dbb68674304

Observation 41bdb891-3502-4a91-92ec-d5d299d47bce · inbound

Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict? cites this paper.

Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict? Great Models Think Alike and this Undermines AI Oversight

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T22:38:10.185108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:38:10.185108Z digest=sha256:5357be08e70b741c465f2e41449adff37475711092f3ceb789616b5552df1c98