Pith. sign in

Paper Citation Record · LEDGER

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools

As of 22 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2508.20410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20410 v3

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:09:57.752216Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T07:24:06.863653Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:53:16.571474Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact8
  • verified fuzzy13
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65aab1c0-7db7-41aa-abff-39dde2965523 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.214481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.214481Z digest=sha256:d388fc7d9f455ab8558bd9311e5071e0ef00cd73b5a7a46f3f69747347263e90

Observation 82ed7fb3-3e55-4177-940b-3c57279370b7 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:00.136771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.228327Z digest=sha256:a186286ab58c13faba937eda16269a1d5f74fa0cf821f94cd6cd21fa9ceddcf0

Observation b0ccbc2a-037a-4b79-b510-5d9d5366ec60 · outbound

This paper cites Bradley and Milton E.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Bradley and Milton E

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.249680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.249680Z digest=sha256:ff87b98b3041d68302da6622f301485d958ebd25b2b404f2d4d74198c33c7e71

Observation 1631b0f7-7cd7-4890-bcdb-45ed640fd8b3 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.259941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.259941Z digest=sha256:670139a3744b859756bd28bfc2a3120d165b6f572be1e7687a39878576402ed4

Observation 815ef57a-394b-4e6a-a9cb-8930cc89e648 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.271533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.271533Z digest=sha256:30c5a877c060980da8ca95b77b26be6b9053800db9682927f3d9559789a6592c

Observation afdfb030-efd2-4bea-9827-9445d805f67d · outbound

This paper cites Christiano, Jan Leike, Tom B.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Christiano, Jan Leike, Tom B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:00.103926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.282278Z digest=sha256:e9557dc486870321897c5c1df7a1f47b81c2954580d96f67c20f1b5f4d99a725

Observation 5164c8bb-d624-499e-90b9-59288ff6734e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.290745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.290745Z digest=sha256:31a18b6ab86118e1d05bf9fbbe747a63ccde2bcc2191e763000382a5deb4684f

Observation 70645d0a-507c-4d93-827f-0c62fb73eff2 · outbound

This paper cites Webthetics: Quantifying webpage aesthetics with deep learning.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Webthetics: Quantifying webpage aesthetics with deep learning

Reference 8

Resolution
verified exact
doi, observed 2026-08-05T15:09:58.246249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.299263Z digest=sha256:85e71b0c39f7a3af6b4ef1a8f5810a5a647b33336c6668e596e0f14b1d11378a

Observation 8052a84d-fe4f-492b-b80a-39e2a558cc6b · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.307467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.307467Z digest=sha256:9066216ccf562ccec471fbfae8a208368aba5c366f652d581845151069df66b0

Observation 7e0ab56e-42f0-47ec-b5e5-71a656ad01dc · outbound

This paper cites How content volume on landing pages influences consumer behavior: Empirical evidence.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools How content volume on landing pages influences consumer behavior: Empirical evidence

Reference 10

Resolution
malformed identifier
doi_truncated, observed 2026-08-05T15:09:58.198420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.325846Z digest=sha256:187a0674f6b62bdf5ae0fce876f8d5dc90fa9e79dc51e5b009f2021c863dc8e0

Observation 1951a0b7-5bd7-4954-8286-58e544bf0fa6 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 11

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T15:09:59.005222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.331145Z digest=sha256:2dbdb080f2128f355a0ee11d2cf6aa5fd87621154e0a8682e477fc0c38e27012

Observation bef0b77a-d46b-4d4b-9eea-904ec3679ed4 · outbound

This paper cites WebCode2M : A real-world dataset for code generation from webpage designs.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools WebCode2M : A real-world dataset for code generation from webpage designs

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:00.076131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.344168Z digest=sha256:5fad95a43880939ae571ef027d6b52bfbdd028f6893b986922ed60beee0ef0c6

Observation d5b66df0-df1a-4725-8add-002146646b19 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Measuring Massive Multitask Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.353378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.353378Z digest=sha256:68a3c26c224d7db66692730bb8efeb2769e3cdd6213eb6736b4213d07d10366b

Observation 10c698a1-371a-4203-831b-a0d9b6bb1e55 · outbound

This paper cites TrueSkill : A Bayesian skill rating system.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools TrueSkill : A Bayesian skill rating system

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:00.037332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.362611Z digest=sha256:80f7c17fd0894180c9f662d6f452bb50f508f780fb651090a1ae7f51007c33df

Observation 805bf1fd-3454-406d-82a7-831e5ee645b0 · outbound

This paper cites CLIPScore : A reference-free evaluation metric for image captioning.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools CLIPScore : A reference-free evaluation metric for image captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.369473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.369473Z digest=sha256:37827f57fc71591cf5728f450ac66c5e12aace8d4daa83f1cc548119f8356703

Observation d096ce4d-9172-4431-ad5a-b9ce72a23dc5 · outbound

This paper cites GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.384528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.384528Z digest=sha256:57daef5ae8d9079085a2f55b5c09298e5eda99ab1ae70f78c3c07a62f2269b20

Observation b33c70c1-dab6-4d35-ad5d-cc14f7848e0a · outbound

This paper cites Jankowski, J.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Jankowski, J

Reference 17

Resolution
verified exact
doi, observed 2026-08-05T15:09:58.131424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.399939Z digest=sha256:b5b5fd083da5c5d50bb89b554370a880c9872ef9679723b241c33ae48a63bc2e

Observation 1af8fce2-5fe6-41c9-8579-99bde460859d · outbound

This paper cites Jayasumana, X.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Jayasumana, X

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:00.007654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.418581Z digest=sha256:eb09081d270c81315f2dc16631faf7e5b4a1785669d66d7f3e25f5442e5aecf2

Observation d6959b45-c4ac-4162-91f4-e7a3ecb6b2bf · outbound

This paper cites Kirstain, A.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Kirstain, A

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.976421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.428490Z digest=sha256:654de490c5c9274f0da32b7ac10b1b1b8cb283d84024f01df97fbb1549e5c7b5

Observation 11f51ac6-fab9-401d-9885-932ddb207ef5 · outbound

This paper cites Assessing dimensions of perceived visual aesthetics of web sites.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Assessing dimensions of perceived visual aesthetics of web sites

Reference 20

Resolution
verified exact
doi, observed 2026-08-05T15:09:58.093488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.435107Z digest=sha256:f2d2aed46e1fea43da2a3355f03036a29452b03342670609f09fdc8d08448788

Observation 4b46df08-7fed-4de0-8419-9a51b9f944c7 · outbound

This paper cites Holistic Evaluation of Text-To-Image Models.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Holistic Evaluation of Text-To-Image Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.445683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.445683Z digest=sha256:28544d80c681c2b75e7250d7c068ac6da9628ba57cfb1ff5cfaca110608d46e1

Observation d375a1df-5c6f-4bfc-b978-cfe80cc22b82 · outbound

This paper cites Lu, Zhe Lin, Hailin Jin, Jianchao Yang, and James Z.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Lu, Zhe Lin, Hailin Jin, Jianchao Yang, and James Z

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.454500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.454500Z digest=sha256:8699374c73adc7b3ae02caa08b930b92a2f1422639107a2e952eb223a37a97b1

Observation cf20d954-4553-47c3-be97-a12dac46b1e1 · outbound

This paper cites WebGen-Bench : Evaluating LLMs on generating interactive and functional websites from scratch, 2025.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools WebGen-Bench : Evaluating LLMs on generating interactive and functional websites from scratch, 2025

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.956095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.473086Z digest=sha256:d10eae57a91f35ddd16a815ad2721a3d722ae862a97aa4a73b53494fd07562ea

Observation ef14ae84-7beb-4282-ba38-330ec87c57ff · outbound

This paper cites FinanceQA : A benchmark for evaluating financial analysis capabilities of large language models, 2025.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools FinanceQA : A benchmark for evaluating financial analysis capabilities of large language models, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.581827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.483190Z digest=sha256:a770c17af6dd5c07c862b8cd29d936b6de5d8ea6047d826c1c1df22c86d25559

Observation 1aa2f16e-85b4-4548-b80f-cd5788622e03 · outbound

This paper cites Moshagen and M.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Moshagen and M

Reference 25

Resolution
verified exact
doi, observed 2026-08-05T15:09:58.052019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.493486Z digest=sha256:029a72215991f95dcc0381feb0676a133d4929b3778e4859820b0c4a5367e524

Observation 5a1b4542-2893-402d-9ec4-9bc09ed90dfe · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 26

Resolution
verified exact
doi, observed 2026-08-05T15:09:58.001466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.498765Z digest=sha256:f2d5003e82e1790809832cf3f5d351cac59aa0aad8398b2c7960381a9e203442

Observation a1a29273-0c60-4209-9cab-56af46a901c0 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 27

Resolution
verified exact
doi, observed 2026-08-05T15:09:57.968857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.508823Z digest=sha256:d2de1058404dd8839e4aa1ae71ab03f387bd9de69296b9dad68a9fd11bd97d3b

Observation 7a3658b6-317a-43be-be2b-509af2ea6a0a · outbound

This paper cites Color compatibility from large datasets.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Color compatibility from large datasets

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.521078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.521078Z digest=sha256:d2f1d7f475cf7396af57468848c5933631b63310318898f8b6c90ac89a1e9b15

Observation 3930f5c6-49c0-4fe0-8bda-23a31511888b · outbound

This paper cites Comparative judgement for assessment.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Comparative judgement for assessment

Reference 29

Resolution
verified exact
doi, observed 2026-08-05T15:09:57.944022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.544692Z digest=sha256:c7588ee0ee829ad809249be0aaa14c157655339fbebb7006e0d3350888a7aa0d

Observation 4a059356-abf0-4b3f-8645-fb68b21e424d · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Learning Transferable Visual Models From Natural Language Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.551803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.551803Z digest=sha256:74ad78ac8255b36ca91b6c4080e5ddc19daccb939a5426d345ede27d43964ef0

Observation 780ed652-e905-4c07-afd6-8047d7423360 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.560815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.560815Z digest=sha256:aec6929878cb41ac2f675430aa9f237f7baca31aeeb2741fc0e511e8fd87dce0

Observation 726bf2ae-d03a-4edd-a1fa-6f783aa1e233 · outbound

This paper cites Robins and J.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Robins and J

Reference 32

Resolution
verified exact
doi, observed 2026-08-05T15:09:57.890926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.567210Z digest=sha256:1783584a5289f360a9a702372e2eac918c0024259754b526b723f8bfa2178c9d

Observation b6e395de-eed0-4f80-bef3-404ac0bf52c9 · outbound

This paper cites Improved techniques for training GANs.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Improved techniques for training GANs

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.528856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.578034Z digest=sha256:637f0df600107dcd63b93f1fc09c9c72b7b5a69a43d41e7905e5a911c7444a41

Observation 00f33793-f036-4738-b5a9-500883cbaa17 · outbound

This paper cites Design2code: Benchmarking multimodal code generation for automated front-end engineering.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Design2code: Benchmarking multimodal code generation for automated front-end engineering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.486372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.587542Z digest=sha256:7bd47d7773f1dd94666e62ac772ede328e6a770f534f9df8289b9eb1c5b65e70

Observation 4bc571d7-7ab6-4154-810c-63e8604c644e · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.600054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.600054Z digest=sha256:70c57a73c1d6d97707a7964f81e54dce181024a268ae3b9576e7754029178523

Observation 7974c531-9c2a-4cf1-8f4c-b6ab4c679265 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.616198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.616198Z digest=sha256:b342d1afdf157d26f3ef24e64004cf3ed0c19a447570c091e760e883e2880c9a

Observation 59b385df-0352-4354-93e7-a1711bde6633 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:09:59.458898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.632179Z digest=sha256:76ef51a5a162d713d16096840743f8880f6653684c37f547662d6a60d338be62

Observation 9e84e18a-f110-41fd-9462-3f84d46f5535 · outbound

This paper cites Whitehouse and A.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Whitehouse and A

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.411804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.660119Z digest=sha256:ecaf3e2ec57198ebbe69e63006293f07b66524541d792e707ca073847c573bc3

Observation 33c07487-797b-403e-b788-78158d38fad2 · outbound

This paper cites The best AI website builders in 2025.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools The best AI website builders in 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.346581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.666058Z digest=sha256:39d0e508f5a4a6707f3892a723b0df478ea668a1cabe5c0114b467e405c72157

Observation 4a486708-4a75-4fc0-b6ee-57bd492e3288 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:09:59.308828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.676210Z digest=sha256:1c0eb66d573b38969c9b31ec41578f5b142d36be0ad4a841110e9592cddf0e2b

Observation 9787ea74-1043-441b-b32d-316a22f13fbc · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:09:59.272758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.685514Z digest=sha256:49a1bb0718d0bd82c102be7a3726b12576bba74f4cc1a32f7dc6bf972ba518b4

Observation 0a914eb3-5578-49f3-9ac9-f97c4e4a3d6c · outbound

This paper cites Xing, Xiaodan Liang, and Zhiqiang Shen.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Xing, Xiaodan Liang, and Zhiqiang Shen

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.218374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.702743Z digest=sha256:de93f2cfb93206e38afb98d9dc0c6523bdf5bac6364f3f063e49f23f1ff8dbb4

Observation 11547017-77de-474a-9487-e17c54314058 · outbound

This paper cites HellaSwag : Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4791--4800, 2019.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools HellaSwag : Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4791--4800, 2019

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.714904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.714904Z digest=sha256:c7ea7d4edc27d50357eeb86a80aca008d55e9dc0bd76e5d01f3cc57f9b510ace

Observation 1ddf4ede-2689-4c44-aa82-7a89064a3b35 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.731423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.731423Z digest=sha256:e21ade0a5d1a52b8ac184dbdbab822e78e2ca2ba4e475c1b32163d06ebdf77c2

Observation 060cab31-313c-40c3-9f4c-a3b8396c4f00 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.740219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.740219Z digest=sha256:45bacf95eafcb5670e2f0d82d3f76da55ab6a953d22f23391359896a01efb9b8

Observation cde96057-7635-47b3-8f39-15b200a6330d · outbound

This paper cites Frontendbench: A benchmark for evaluating LLMs on front-end development via automatic evaluation, 2025.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Frontendbench: A benchmark for evaluating LLMs on front-end development via automatic evaluation, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.186704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.752216Z digest=sha256:79d7027465eabf4737a2bf29430f038d36b114cdfbf33f00bfb7dd408926ce32

Pith citing papers

Observation 7f67039e-6fbf-438a-ab82-ba6c45131e25 · inbound

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation cites this paper.

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.572734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T07:24:06.863653Z digest=sha256:88de53f8895e5251d1bb4456f7c13d9e48e9b9ce7066cc8d427b165ff07e7aa5