Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:54.867626Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.15773.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:54.867626Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 873e7234-a520-4a73-b01c-7b87849f185a · outbound
Supernova: Achieving More with Less in Transformer Architectures Attention is all you need.Advances in neural information processing systems, 30, 2017
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e15f32-f3eb-482a-bfb4-baa52963afdf · outbound
Supernova: Achieving More with Less in Transformer Architectures Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 134bb493-ef2b-4749-9a81-67f508589e2a · outbound
Supernova: Achieving More with Less in Transformer Architectures BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 035b3f49-bd53-4177-9e9f-d631df31a567 · outbound
Supernova: Achieving More with Less in Transformer Architectures RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81d75129-1e8e-43d0-aa96-c1c6ad580f70 · outbound
Supernova: Achieving More with Less in Transformer Architectures GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d11eba53-b563-4566-8045-024fde1b974e · outbound
Supernova: Achieving More with Less in Transformer Architectures Fast Transformer Decoding: One Write-Head is All You Need
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 336e5187-496b-4861-b556-f88c7814ee84 · outbound
Supernova: Achieving More with Less in Transformer Architectures GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1341b25f-ce54-429d-8299-968cc77b1380 · outbound
Supernova: Achieving More with Less in Transformer Architectures Layer Normalization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b43304-a96b-4fc6-9500-85abdb6a70c4 · outbound
Supernova: Achieving More with Less in Transformer Architectures Root mean square layer normalization.Advances in Neural Information Processing Systems, 32, 2019
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20b91860-2f11-494c-9ab1-377d2724ef77 · outbound
Supernova: Achieving More with Less in Transformer Architectures Language modeling with gated convolutional networks.International Conference on Machine Learning , pages 933–941, 2017
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a1dad2b3-adf3-42c5-b285-c1e26204d926 · outbound
Supernova: Achieving More with Less in Transformer Architectures GLU Variants Improve Transformer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 212bc820-8b48-4381-919b-7ef2f9913709 · outbound
Supernova: Achieving More with Less in Transformer Architectures PaLM: Scaling Language Modeling with Pathways
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2aa7d57-4fc9-4169-971a-822793485027 · outbound
Supernova: Achieving More with Less in Transformer Architectures Neural machine translation of rare words with subword units
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a12f0412-698e-4ae9-8ab8-436d5878f182 · outbound
Supernova: Achieving More with Less in Transformer Architectures Byte pair encoding is suboptimal for language model pre- training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 34d8732a-d737-4ae6-b7e8-a280f3e7dabb · outbound
Supernova: Achieving More with Less in Transformer Architectures Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3275c2c9-fecb-43dc-9e70-da1cf701be01 · outbound
Supernova: Achieving More with Less in Transformer Architectures Textbooks Are All You Need
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c786c9-768c-49f1-a0b8-4b6e3057179d · outbound
Supernova: Achieving More with Less in Transformer Architectures Stable LM 2 1.6B Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d00e7bc-1714-42b2-8332-69ea21f8da41 · outbound
Supernova: Achieving More with Less in Transformer Architectures Gemma: Open Models Based on Gemini Research and Technology
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5df4b5fa-9275-4f5e-aa9b-b4a14bf91e39 · outbound
Supernova: Achieving More with Less in Transformer Architectures Scaling Laws for Neural Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09a9d5ea-e062-4abe-9e1b-5504efcc3ee2 · outbound
Supernova: Achieving More with Less in Transformer Architectures Training Compute-Optimal Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ed90807-a9b7-4aef-880f-a1eb6ab2360b · outbound
Supernova: Achieving More with Less in Transformer Architectures Scaling Data-Constrained Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad5398d7-1a52-4244-bdf8-a35b0ecc6dd3 · outbound
Supernova: Achieving More with Less in Transformer Architectures LLaMA: Open and Efficient Foundation Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.