Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:09:57.752216Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2508.20410.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:09:57.752216Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T07:24:06.863653Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T08:53:16.571474Z
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 65aab1c0-7db7-41aa-abff-39dde2965523 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ed7fb3-3e55-4177-940b-3c57279370b7 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b0ccbc2a-037a-4b79-b510-5d9d5366ec60 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Bradley and Milton E
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1631b0f7-7cd7-4890-bcdb-45ed640fd8b3 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 815ef57a-394b-4e6a-a9cb-8930cc89e648 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afdfb030-efd2-4bea-9827-9445d805f67d · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Christiano, Jan Leike, Tom B
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5164c8bb-d624-499e-90b9-59288ff6734e · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70645d0a-507c-4d93-827f-0c62fb73eff2 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Webthetics: Quantifying webpage aesthetics with deep learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8052a84d-fe4f-492b-b80a-39e2a558cc6b · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e0ab56e-42f0-47ec-b5e5-71a656ad01dc · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools How content volume on landing pages influences consumer behavior: Empirical evidence
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1951a0b7-5bd7-4954-8286-58e544bf0fa6 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bef0b77a-d46b-4d4b-9eea-904ec3679ed4 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools WebCode2M : A real-world dataset for code generation from webpage designs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d5b66df0-df1a-4725-8add-002146646b19 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Measuring Massive Multitask Language Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c698a1-371a-4203-831b-a0d9b6bb1e55 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools TrueSkill : A Bayesian skill rating system
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 805bf1fd-3454-406d-82a7-831e5ee645b0 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools CLIPScore : A reference-free evaluation metric for image captioning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d096ce4d-9172-4431-ad5a-b9ce72a23dc5 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b33c70c1-dab6-4d35-ad5d-cc14f7848e0a · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Jankowski, J
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1af8fce2-5fe6-41c9-8579-99bde460859d · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Jayasumana, X
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d6959b45-c4ac-4162-91f4-e7a3ecb6b2bf · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Kirstain, A
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 11f51ac6-fab9-401d-9885-932ddb207ef5 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Assessing dimensions of perceived visual aesthetics of web sites
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4b46df08-7fed-4de0-8419-9a51b9f944c7 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Holistic Evaluation of Text-To-Image Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d375a1df-5c6f-4bfc-b978-cfe80cc22b82 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Lu, Zhe Lin, Hailin Jin, Jianchao Yang, and James Z
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf20d954-4553-47c3-be97-a12dac46b1e1 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools WebGen-Bench : Evaluating LLMs on generating interactive and functional websites from scratch, 2025
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ef14ae84-7beb-4282-ba38-330ec87c57ff · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools FinanceQA : A benchmark for evaluating financial analysis capabilities of large language models, 2025
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1aa2f16e-85b4-4548-b80f-cd5788622e03 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Moshagen and M
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5a1b4542-2893-402d-9ec4-9bc09ed90dfe · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a1a29273-0c60-4209-9cab-56af46a901c0 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7a3658b6-317a-43be-be2b-509af2ea6a0a · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Color compatibility from large datasets
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3930f5c6-49c0-4fe0-8bda-23a31511888b · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Comparative judgement for assessment
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4a059356-abf0-4b3f-8645-fb68b21e424d · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Learning Transferable Visual Models From Natural Language Supervision
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 780ed652-e905-4c07-afd6-8047d7423360 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 726bf2ae-d03a-4edd-a1fa-6f783aa1e233 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Robins and J
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b6e395de-eed0-4f80-bef3-404ac0bf52c9 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Improved techniques for training GANs
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 00f33793-f036-4738-b5a9-500883cbaa17 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Design2code: Benchmarking multimodal code generation for automated front-end engineering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4bc571d7-7ab6-4154-810c-63e8604c644e · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7974c531-9c2a-4cf1-8f4c-b6ab4c679265 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59b385df-0352-4354-93e7-a1711bde6633 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9e84e18a-f110-41fd-9462-3f84d46f5535 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Whitehouse and A
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 33c07487-797b-403e-b788-78158d38fad2 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools The best AI website builders in 2025
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4a486708-4a75-4fc0-b6ee-57bd492e3288 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9787ea74-1043-441b-b32d-316a22f13fbc · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0a914eb3-5578-49f3-9ac9-f97c4e4a3d6c · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Xing, Xiaodan Liang, and Zhiqiang Shen
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 11547017-77de-474a-9487-e17c54314058 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools HellaSwag : Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4791--4800, 2019
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ddf4ede-2689-4c44-aa82-7a89064a3b35 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 060cab31-313c-40c3-9f4c-a3b8396c4f00 · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cde96057-7635-47b3-8f39-15b200a6330d · outbound
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Frontendbench: A benchmark for evaluating LLMs on front-end development via automatic evaluation, 2025
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7f67039e-6fbf-438a-ab82-ba6c45131e25 · inbound
Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.