Pith. sign in

Paper Citation Record · LEDGER

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

As of 20 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 5 inbound Pith citation observations for arXiv:2507.03865.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03865 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:08:09.488803Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:15:07.167572Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8025c31f-e790-499d-bb46-40311bd8d76e · outbound

This paper cites write newline.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.520040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.520040Z digest=sha256:1d44af2448e9be34b939f5ac6b073691db0c9897b54d5e0249aae13878b9f91f

Observation 3271328d-b004-4830-869d-45401ba497cd · outbound

This paper cites Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.144777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.540763Z digest=sha256:4059846c231e8907c4bd3fb073c73f12172ea0c76e1894b46a8f53161b8f94c5

Observation 4861d466-5830-4fd9-b861-0622edabd0d0 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Longbench: A bilingual, multitask benchmark for long context understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.134965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.575951Z digest=sha256:0f7652a492c60afce26f22244ef2869a1f81d0256633765a1a7e90e9ae698ede

Observation a3febf6d-418c-4022-a485-f1c4592d3c68 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Piqa: Reasoning about physical commonsense in natural language

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.665742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.665742Z digest=sha256:2464d7acaace0853d9296ff44d5bc0919d4c12dc2c2c73c56a52553019e12e65

Observation 7b46542b-851f-463d-8eef-d25b67658055 · outbound

This paper cites Spectral filters, dark signals, and attention sinks.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Spectral filters, dark signals, and attention sinks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.694850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.694850Z digest=sha256:ed71f7a1957425d236a0a257af3f97b115b88e9a0a8eefa88c5930c0b34639bf

Observation 7e3cbe33-ce91-4136-8db3-9c084ff0c855 · outbound

This paper cites SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.728594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.728594Z digest=sha256:ec4e72384a88e24e6ad4339973b2d64eac63a185766498adab8c0827f780b085

Observation 3871d836-3f1b-46ac-99f7-e2d643164297 · outbound

This paper cites Ee-llm: Large-scale training and inference of early-exit large language models with 3d parallelism.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Ee-llm: Large-scale training and inference of early-exit large language models with 3d parallelism

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.119715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.784271Z digest=sha256:573e1f4efddcc1361ee0bd810025cb21c12ab578a0b5d0df2795bf6378bee03f

Observation 6a2e375d-4bd9-46bc-b2f7-ceee1f13df11 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.826197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.826197Z digest=sha256:34016f081d03243fa7427fced003a80430fa85380e7f72182131d7a728eef2e1

Observation 8fe2c9d0-f948-465b-942a-3381ee8bc030 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.884228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.884228Z digest=sha256:113255bdfcb5379a1182c7bfcb9cd12bbfccc14e2c5baa428462c470132e11e1

Observation 22d001fd-06b5-44b3-9188-35a713a1b314 · outbound

This paper cites The Llama 3 Herd of Models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:06.940790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:06.940790Z digest=sha256:0a47c2b04d3464b6ea7a7d6200c69b23cc85a844b3087c3dea7ed2d9da529c73

Observation cf401c8b-e256-4a47-9367-6c7a4db9f037 · outbound

This paper cites Layerskip: Enabling early exit inference and self-speculative decoding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Layerskip: Enabling early exit inference and self-speculative decoding

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.110224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:06.986186Z digest=sha256:7af39c78b9f9f5fe854fad856d1fbe0fdee406101df3dc6bd4add4f7579e8583

Observation 1c73407b-f2ef-4a55-b6bb-09fa8cbfb7b4 · outbound

This paper cites When attention sink emerges in language models: An empirical view.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference When attention sink emerges in language models: An empirical view

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.099806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.062613Z digest=sha256:eac52b69c2ca52e55d6ed8d558169e18639e6c2314eea5ceeb44c3eace5e81b9

Observation 870b8fee-af36-484a-8381-11b1f0f59ac2 · outbound

This paper cites J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.136253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.136253Z digest=sha256:5004f629abc51ee85e0e2c5a77cf517a0d26451bfbca21a89d5adba393a75b82

Observation 3133fe33-7106-4036-8df3-ea59d7c6b30b · outbound

This paper cites Mistral 7B.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.195766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.195766Z digest=sha256:2a64c4ce0ec0417e7d876f6a50d5b629cb709b64407ec079157b0453be2372fa

Observation 52bdaca1-fa16-4b5d-b930-9629ec47723e · outbound

This paper cites Mixtral of Experts.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mixtral of Experts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.249338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.249338Z digest=sha256:a366bd72b500f8aff7a6fd9689b21a7f837a4d379657fcdce6bfa978a89b4393

Observation c5ab7db0-7fb5-4b79-a78f-e3774e77afb1 · outbound

This paper cites D-llm: A token adaptive computing resource allocation strategy for large language models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference D-llm: A token adaptive computing resource allocation strategy for large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.082034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.288318Z digest=sha256:e352b9b1280d83ea7575ff4f7ff85e388c3b0eacb2a72c4c90a5b63e7014afc4

Observation b92bd581-f141-4ed6-955b-46c974e5ccc7 · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.404789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.404789Z digest=sha256:9647ed4d4d1a5971e924a498ed4c121f446c683edb7429067ed86e5c40094edf

Observation d0c73403-f92f-4756-93ab-f13daf0b7375 · outbound

This paper cites B io M istral: A collection of open-source pretrained large language models for medical domains.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference B io M istral: A collection of open-source pretrained large language models for medical domains

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.071971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.480577Z digest=sha256:622e453a5bf980c0f406711f692fdb048f0bfef9bf1a37f509b16e6eb452ebcd

Observation 00e02adc-9716-4445-82a2-57fffb6784c8 · outbound

This paper cites Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.633297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.633297Z digest=sha256:491a5bdd64a09ae1f9ed4e6900e34d072c79b6330ecb140f78cce03f9cf5ab91

Observation ebc7bd35-3c72-41ea-8248-cfa28d1b62d4 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.685132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.685132Z digest=sha256:860c7fc55faf8e6e55779d9c433959e213dd49524d9f0ea9c4c1cbb631638be4

Observation b5ce3c25-d6bf-44cd-8c55-7a4c02fc4a5f · outbound

This paper cites Pointer sentinel mixture models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Pointer sentinel mixture models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.061210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.752189Z digest=sha256:c8188fb700e4f3381e661afbae96bfc67d787b02380f258a40fb5520294f5d0b

Observation 7e696759-321d-482e-a7b2-db9808987c1a · outbound

This paper cites Using an llm to help with code understanding.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Using an llm to help with code understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.050576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:07.806617Z digest=sha256:c437c98bea0c6eb0444bc9a9e8c962f4dd0b5ab5251260ab240b70ac6bd6822f

Observation e3f9296e-de56-4039-9696-d66615f753ec · outbound

This paper cites an unresolved cited work.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.846759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.846759Z digest=sha256:70197dd30710464d2f48c9702af765d17bc38a879f3da1281347e91a0559ae56

Observation 57f4bdd9-3066-46f7-b105-9bcddaf9047a · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.957269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.957269Z digest=sha256:67349723ccf41d9a46c9f1621c18ed2ac3209a0f6957ee155260a1a4854e6363

Observation e5224ed9-8b43-4c75-ac39-c02ce6f43d03 · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference L., Bhagavatula, C., and Choi, Y

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:07.997272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:07.997272Z digest=sha256:86ef400c66d490f57fe53bea2003f77f76b204d72e10a2e2e4d7b0655c37edc1

Observation 67011f70-86fa-4438-b248-fe500a267bcb · outbound

This paper cites Q., Tay, Y., and Metzler, D.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Q., Tay, Y., and Metzler, D

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:11.027343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.188012Z digest=sha256:62b2501d5356376e938fb6ee94d28782f6c4d38193c2677bec3c69f7c9a339da

Observation 06e44b8d-493f-4b07-9890-2f66f876fd32 · outbound

This paper cites A deeper look at depth pruning of LLMs.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference A deeper look at depth pruning of LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.244465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.244465Z digest=sha256:65bd77a94a9d18d6fc387fa00e5d8567ee28db9f1e6034e20a90fbd0a83d9503

Observation e2bb6b72-cc2a-47c0-9b84-e450ce09f459 · outbound

This paper cites Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.298473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.298473Z digest=sha256:949aba23ebb790488fb984aaf43284f205a42541d8eef6ded4aec81b477b3566

Observation 32042cfd-8cb9-4108-b8ba-e4bf8c82d96f · outbound

This paper cites Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.915706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.393078Z digest=sha256:0175767ec525e655e55aa9e2136af06dfa0d31e35327fabd653587ca3bb89691

Observation c8970a42-2ccc-4f37-8dd1-ff8b0dd3fc6e · outbound

This paper cites Z., and Liu, Z.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Z., and Liu, Z

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.906428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.446045Z digest=sha256:078771d9090e28539f6b735adcc376698808fc1eff5300511001492fbc314d6c

Observation 6fc4520c-285e-49a0-b59a-02cabd591125 · outbound

This paper cites Razorattention: Efficient kv cache compression through retrieval heads.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Razorattention: Efficient kv cache compression through retrieval heads

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.896623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.512199Z digest=sha256:3d8b376d5536295417616729a4098ad264a9e7d4743421b3fcb7d852362c0418

Observation 3846996a-fad5-4ef1-9520-6dc86ad22fcc · outbound

This paper cites J., Ting, D.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference J., Ting, D

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.740169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.610825Z digest=sha256:fd8106ba1e5af572436de7bf925f71f2d5842b6502ce271c85f01eddafec31db

Observation 4533b7cf-cd02-4b68-9921-77c3bb68dc7f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.722523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.722523Z digest=sha256:b385e0f46e1681d3f768dfc7c25d35e5169f99dc0283f604c875bbcd1912aa8e

Observation bd2fa558-17d1-4cef-9907-f074a86bf813 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference BloombergGPT: A Large Language Model for Finance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:08.837488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:08.837488Z digest=sha256:b7b371ab77daf446f15f068857829d0e0b7216103f439a51827607da16475d5e

Observation 2877ea30-5aea-4eb7-8810-eddfceecbbbc · outbound

This paper cites T., Peng, R., Wu, Q., and Wang, C.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference T., Peng, R., Wu, Q., and Wang, C

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.551322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:08.895926Z digest=sha256:53f901e19819822ac8af80ef66e886a036319196005bcba3a3b18b4536a2faa5

Observation 13964cca-0020-4437-8b2f-7895d5678333 · outbound

This paper cites Duoattention: Efficient long-context llm inference with retrieval and streaming heads.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Duoattention: Efficient long-context llm inference with retrieval and streaming heads

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.339924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.003068Z digest=sha256:2038f051f4908a1ca62866f6c287d92fb68887b70c60fc5d5f320b3b7912b2f8

Observation 8527fb09-f508-40c1-a700-609a7a9bc288 · outbound

This paper cites Efficient streaming language models with attention sinks.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Efficient streaming language models with attention sinks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:10.169486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.106850Z digest=sha256:46ab9babc40a4b168349470a67ec4c2b2f2bd8a8d338d7e7b5d442a788ec16de

Observation 3ee6e21e-26f5-49f1-a821-95061880a634 · outbound

This paper cites an unresolved cited work.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:08:10.029542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.248931Z digest=sha256:8cbdc96df2a5c8901865fce11952e76cc86dd180d9eae7c92ce0b5cd1625b4ff

Observation 7f2a254e-64b3-4b64-a42b-e0c7054921df · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 4791--4800, 2019.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 4791--4800, 2019

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:08:09.806985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T20:08:09.372102Z digest=sha256:cbfbbfcc19e6ddf2677ed530e305b5d2d9d08a5e5d753de18166332a15ac3f07

Observation ae04f335-09c4-4713-9b2e-f61df35db73e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:08:09.488803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:08:09.488803Z digest=sha256:e806ed01966b471f6e606394f7278e31c492616cd73a4864eff11df4207a1afb

Pith citing papers

Observation 4040624e-1a52-449c-ad62-3e59d3d9a76d · inbound

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers cites this paper.

DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:15:07.167572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:15:07.167572Z digest=sha256:7d2dd42eaedc6d2db4a4b9ef5f06a57fce7ef7a1041f96f26d2abcbd8a049c55

Observation 1b256f7d-5bfb-4f70-a411-d8e74ae2a03f · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:47:37.280027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T08:47:29.236561Z digest=sha256:5a98f64f1afe9a1a51ffb868a1a59870033011134fcf1274ccfb9f5da55defc0

Observation e4d1dbc0-a3f4-46bd-87cd-26436f2198d6 · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T05:50:25.027494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:50:25.027494Z digest=sha256:ce9e912a8cfac46948a15428f6764fbe0b49d7941cc1d9014528a86bddc7b36d

Observation df018183-1fe6-4e00-831b-dc59de7a9fd1 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:32:46.707032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:ac05a2e0e60f405c8f95df5c1d814e6046389cb73520f30bcf49278ea2755744

Observation e37ca751-0043-4a4b-9ad3-43f1489a797d · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:47.283100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:2762caf82e4a4a52efdb47e21a72eebdf06a67bec655b77744756e5379772873