Pith. sign in

Paper Citation Record · LEDGER

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models

As of 14 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2505.20645.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20645 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:55:26.474627Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:21:09.104594Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T04:21:09.502392Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3678a104-4fb6-4ba7-b583-62d3c55e9982 · outbound

This paper cites online" 'onlinestring :=.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:23.031214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:23.031214Z digest=sha256:f9d7be3327385477ec6def07ae9c1a127feea69c090e8cc335aa4858d099226e

Observation 9a2f5bd2-6d8e-49ac-a2a9-73bc4c4f41c6 · outbound

This paper cites write newline.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:23.164159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:23.164159Z digest=sha256:f665294fffc54482dbaf8edc992ba871a9fd86e75cee6220692a16cbf8585080

Observation 3d64b0ce-6a56-4875-93c3-e0152a72c205 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:23.301304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:23.301304Z digest=sha256:2d17180ec20660b9525c0e0b9de6b921ad953608f04f27e3ebad9e948f32c2e5

Observation 7fed5dda-db73-4f3a-b954-95000fee73a0 · outbound

This paper cites Steering Large Language Model Activations in Sparse Spaces.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Steering Large Language Model Activations in Sparse Spaces

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:23.396550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:23.396550Z digest=sha256:536f195a6bf933946017045a5468c1cfa025c4b6a9713fdd076cf2cab5e040ab

Observation d1e4f3fd-8550-4f8f-8ff7-57771923261d · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:28.967099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:23.469895Z digest=sha256:c721d04ffd39c3497777c8dda2db3decb2cc9f75118f81ac075b80631af54806

Observation f4a29e80-2618-4887-a8af-bbaff5786fda · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:28.779620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:23.544602Z digest=sha256:3211d7b7a5ff7750e63134bddd2206cfb13d2c1319b08cede59738e7b0138763

Observation 03509357-a1a6-4528-88c9-eafe6a1bcaf0 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:23.646609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:23.646609Z digest=sha256:c84535422d54c8d910ac7e0d53738e540d1f9fcae02b4875fda2b537927079a3

Observation 6b99d4f3-d131-4f29-8441-94681a182eee · outbound

This paper cites The Oscars of AI Theater: A Survey on Role-Playing with Language Models.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models The Oscars of AI Theater: A Survey on Role-Playing with Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:23.756147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:23.756147Z digest=sha256:1ccf5b5a57159d034577e833f9a80d24ab79142e53141cc60dc72e2d0ffe79cf

Observation f0844eaf-74da-4531-84f5-26950888537f · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:23.877640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:23.877640Z digest=sha256:e9bdc194cf31f100e5d3d719145d2ae7f56ceb8a228c6803dd398b481cc44939

Observation 9189b25b-60b9-4cc0-b65c-cbecb2670ef6 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:28.571705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:23.959641Z digest=sha256:3846caeea27672ae2e357d277a01b150eb12f677045a7cb6f2dd060bc2c237ae

Observation 08ccbc95-eb83-45ba-835a-c09817365ff4 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:24.021433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:24.021433Z digest=sha256:9b2214a32d7eba8409c232fadc0144eca6e3978be10764ab91cf35c44b2b83a3

Observation f29105e1-6cb2-4ccb-ad53-eef96b8fe982 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:28.402652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:24.105466Z digest=sha256:a445362193c50fe03e4e836059a240eb4d34ad8f3d91ee3e96bcf7e52b3b40a2

Observation 806d745c-ddd5-408f-82f9-1597793be352 · outbound

This paper cites BERTopic: Neural topic modeling with a class-based TF-IDF procedure.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models BERTopic: Neural topic modeling with a class-based TF-IDF procedure

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:24.177954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:24.177954Z digest=sha256:921f385d59d7be3851f555c3e023d43a1aa20ddc6387f8920576736e9686ad3c

Observation 2bbf2f84-1d77-4032-956e-52218ecdd21e · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:24.215136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:24.215136Z digest=sha256:b5d6572c8ae581ee23c188922d800326b7ee00013ed8ef4508366d60cab855d5

Observation bd7f1696-27dc-478c-a417-f14fa3e506e5 · outbound

This paper cites a m \"a l \.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models a m \"a l \

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:28.260944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:24.265170Z digest=sha256:e4ef062e82a33f7e6b7d4d1b0d8077af01942e5310ef8a2aa52753b69e6a9863

Observation c524bf74-e85a-49cd-9876-4a8308cdcc8a · outbound

This paper cites From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:24.305287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:24.305287Z digest=sha256:6d84f9769d1aac72bfb9e9c53d7026744ea2d534a4bd094804ad073b0390b42d

Observation d5c68d77-fab0-40a9-93df-378883fa50d9 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:24.371150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:24.371150Z digest=sha256:b194a748b9a033dd8b21e5ec7f23eb69b9b2bf113d0b5336f7480f69ee285729

Observation 6829a3cd-4bac-4e58-a786-34465323e727 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 18

Resolution
verified exact
doi, observed 2026-08-07T13:55:26.696008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:24.449316Z digest=sha256:a3f7f4ae075f0eacf10066e2b83f46ea9e5ee37ae615d9ef982451b7eb0ad346

Observation 341012a3-0db8-46d9-b2bb-bd601dab9778 · outbound

This paper cites Aligning ai with shared human values.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Aligning ai with shared human values

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:28.164053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:24.514927Z digest=sha256:5657ec5a3732dae85c8884708fbf49e434243cf92125a028672584197c1387a4

Observation 0201ebbc-74d1-4525-8e8c-ce6b79439072 · outbound

This paper cites Smith, and Hannaneh Hajishirzi.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Smith, and Hannaneh Hajishirzi

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:55:28.021854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:24.591601Z digest=sha256:07c29222cf17bf34dfda6f0c90154602f3f21a548f8279b75ca39d5ffbfba44d

Observation 142e94a9-8af7-4201-9cf1-f593f5170a7f · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:24.666573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:24.666573Z digest=sha256:cd491fcbfd31b1a2329eb297288b03c6f8e0c6f7e4d3c15b29f1c86c6ded6820

Observation 320203d5-3a3e-43a3-bdeb-1b2f68262e7b · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:24.764658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:24.764658Z digest=sha256:76ffdc313117dc7a61fdea85888bd431adf91678a1aa917d3567eb6e4b4865c0

Observation 5a2905ac-aa31-4006-9fef-2a72a38038ad · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:27.914835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:24.838996Z digest=sha256:2e760582f6eefa397a2891851d7e4bc741a89a5771854104ef959c9f6dddcc96

Observation 5172d1ae-6a34-446f-acd1-f875e135d0b8 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:24.895883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:24.895883Z digest=sha256:54ded8d53032f9c47db608e5511db3e03f8c7fb4a5c74c4c29e994094264314c

Observation f94c4197-7cfe-4008-8aa4-4fcea416ffd9 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:27.796003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:24.981143Z digest=sha256:2f7d5f347df0350dfd2c8063ab21e1130fafd3a3e50d344ff05f4c96cf3ae6da

Observation 10feda8b-dfb9-4b64-8be4-60d2c58ca271 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.063989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.063989Z digest=sha256:c6a6ddedef6c2457b6cdc978644ec73fbde7db59124626830221234ef32de356

Observation e153c7aa-c5c8-464e-9078-e319e3e12d75 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.167622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.167622Z digest=sha256:e03432f656f23dc31684fd8fdb4f14e33e0f7d81965843d40dd88540c7c94b07

Observation ded4939f-32f4-4a99-b7e4-6b20b583700e · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.210926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.210926Z digest=sha256:edcc24cb2064df3b3b4bfe82cb67ab11342705e851e92817835596ab229be6f7

Observation 3d982a9b-72ab-4994-9114-23c61d3905ee · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.275767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.275767Z digest=sha256:26922c1f43f08a39705be40f8e821a89ab702b4a9d320c72ac44a4a11c539209

Observation bbd91e97-fe96-4c32-8a49-1cc55e6f992f · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.347527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.347527Z digest=sha256:3297d24d2aa51db07a1efd34cffbcd33e4db0fc185a85ea96f83e1b7659fcac5

Observation bce8bca1-0acd-4c7e-973a-5fff009596c1 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.415397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.415397Z digest=sha256:3523476898a3db4a0d57c08e6ea9bc71848c75bca19d0c1fa414800c951610f0

Observation 77aa0b0c-fd8b-4f74-b852-c23db2761c7c · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.490965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.490965Z digest=sha256:700560b7c99faafefb5a498159beb9bd48539d27eb3d522de811d9c4adb4373c

Observation 353ef87f-509a-4f8b-864f-aa2b54501e17 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.537083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.537083Z digest=sha256:774cbc5c38b5cc931eb7cffdfe7528872e2d0f4e7516070205b97ac6060be716

Observation 6f5adc31-3497-48aa-9acc-c271fc58ad3c · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.607345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.607345Z digest=sha256:1ebc841dda984d485520c4d7a13b998afff6e89a166b00e406a27412018fbcfe

Observation e52aa194-98a4-437f-a161-ffcdd34fb001 · outbound

This paper cites Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.697100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.697100Z digest=sha256:aa11a056de2b705506646ab12923cadd5cfcb8e0858d259b518a7405ca541621

Observation 060c3749-cb58-417e-ab6d-95e70ad157ec · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.774971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.774971Z digest=sha256:9f9b90b86824f63d2ca2c15f03a2b86dc4ad08717eb825196082a8ccec87f456

Observation a2682378-3c5c-4aae-b735-c460e0eb1941 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.821115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.821115Z digest=sha256:af32df95e54569043bc2cf9a5ebe60090c2e65d0f696c76284f93d44a7516cc2

Observation 3b542f9a-63ee-458e-804a-5d4cb0dda103 · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.853785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.853785Z digest=sha256:eb4b39ece45a8434f8081b8ba0197a0d6c6d4d1ce6c05dcf0dbc479f89db981d

Observation 7f2e6fcc-87c9-4235-8cf4-4b186d2b7140 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:25.908255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:25.908255Z digest=sha256:e9029eb8e4d9dbe41f59a91172c497b0ac2a3d4d1a89a922a957adb581197bc6

Observation 9c1a650e-6a35-4559-a5cd-2d28fe02dd87 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:27.657272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:26.012147Z digest=sha256:5f338b0ea0543bcb6340f843d932b1895f10d2841df4e3cd8e7dad0e8c5fa47d

Observation cd076df4-4dea-48fc-95ff-a02506e5f376 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:26.074596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:26.074596Z digest=sha256:da4f362d68c58522d482edf39b5569e5323ba9161e321e884d0f718f54acefa9

Observation 96d78548-34e7-49c8-b13e-c72d4c514d40 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:26.156757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:26.156757Z digest=sha256:90cfafb89849246717dd0672f534c223522a52a1b905bd3e8d71a9b2711df283

Observation b4dfd2be-3187-4604-9d96-bade2c08ff9d · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:26.240689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:26.240689Z digest=sha256:25741064558ff34422fea5274fa4f20f971eb25217dd2256438cd7e4fe19088a

Observation 87026bb3-9943-4c81-b678-b9266d4d4563 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:27.529932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:26.293417Z digest=sha256:3c2920180af7892af838f80515d8112541923e324e13c2e8409e39e2f54cf675

Observation bc408312-14ce-4608-afd9-4918149dbf08 · outbound

This paper cites an unresolved cited work.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:55:27.402408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:55:26.380658Z digest=sha256:4870df832799dc0ea189786ccf12f112d81469069c9c743a8b6411bbf2bb905a

Observation 46d49f00-0b25-4345-9b97-5c13ef782fda · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:26.474627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:55:26.474627Z digest=sha256:857592523e020e216fcfd89a014a3d4e3ed803c48537f5ede14ea8385e2cb487

Pith citing papers

Observation 84b27609-a3c0-4f32-8be9-6cf62598460f · inbound

Divergent Response Modes in Frontier Language Models Under Steering Pressure cites this paper.

Divergent Response Modes in Frontier Language Models Under Steering Pressure STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T04:21:09.511371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-10T04:21:09.104594Z digest=sha256:8735ec0fa91a9a86f40ed6a25da4ba4b64855da148a09866dd8a939053297564