Pith. sign in

Paper Citation Record · LEDGER

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools

As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2508.20410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20410 v3

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:09:57.752216Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T07:24:06.863653Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:53:16.571474Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact8
  • verified fuzzy13
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65aab1c0-7db7-41aa-abff-39dde2965523 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.214481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.214481Z digest=sha256:9fd69d2c5dd5c5d9796992eff82eb5fb13ef5369af72bdff4ae1920093553dff

Observation 82ed7fb3-3e55-4177-940b-3c57279370b7 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:10:00.136771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.228327Z digest=sha256:c0f59ca3183708742c4c52785f720a308aff665e46dae35c918a25d71c299c5e

Observation b0ccbc2a-037a-4b79-b510-5d9d5366ec60 · outbound

This paper cites Bradley and Milton E.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Bradley and Milton E

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.249680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.249680Z digest=sha256:1de2aa987c9312a64a7cda6c3dc9e2303ea3ec445b4de374787312c587d8a960

Observation 1631b0f7-7cd7-4890-bcdb-45ed640fd8b3 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.259941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.259941Z digest=sha256:d564584d6922163c674cb1d02ab7b4dfe66038cc72d57301b1dcc55c6b0aa570

Observation 815ef57a-394b-4e6a-a9cb-8930cc89e648 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.271533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.271533Z digest=sha256:3dd08dcde1849b473d940098a4d5ecaaa3cb8d5144ca19c1a2471835651da18c

Observation afdfb030-efd2-4bea-9827-9445d805f67d · outbound

This paper cites Christiano, Jan Leike, Tom B.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Christiano, Jan Leike, Tom B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:00.103926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.282278Z digest=sha256:556d2ccac6f4521d9075f4a996d46aa24d75b3cefafc8b4d766aca52c109366f

Observation 5164c8bb-d624-499e-90b9-59288ff6734e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.290745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.290745Z digest=sha256:a3caee58afd553c6760a1a28c59d3d8f2392fee887c6eb725fb5e65943c51650

Observation 70645d0a-507c-4d93-827f-0c62fb73eff2 · outbound

This paper cites Webthetics: Quantifying webpage aesthetics with deep learning.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Webthetics: Quantifying webpage aesthetics with deep learning

Reference 8

Resolution
verified exact
doi, observed 2026-08-05T15:09:58.246249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.299263Z digest=sha256:da166c1bfd93796a5d4d12c136406e355efb88332ec21e7f52e80b650f58657d

Observation 8052a84d-fe4f-492b-b80a-39e2a558cc6b · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.307467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.307467Z digest=sha256:9b57f35d2a4923dcba01e9edde00935b2c450c52e3104ea6b2e5d056c71a4838

Observation 7e0ab56e-42f0-47ec-b5e5-71a656ad01dc · outbound

This paper cites How content volume on landing pages influences consumer behavior: Empirical evidence.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools How content volume on landing pages influences consumer behavior: Empirical evidence

Reference 10

Resolution
malformed identifier
doi_truncated, observed 2026-08-05T15:09:58.198420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.325846Z digest=sha256:95fa5463ca31504d31d7cbf6b9fc4d78ec6e91af7e619e1cf9908adf6eada4ab

Observation 1951a0b7-5bd7-4954-8286-58e544bf0fa6 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 11

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T15:09:59.005222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.331145Z digest=sha256:785f07b192cd298195c189897465b273186933e15819e86f1caa3740ae6afefa

Observation bef0b77a-d46b-4d4b-9eea-904ec3679ed4 · outbound

This paper cites WebCode2M : A real-world dataset for code generation from webpage designs.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools WebCode2M : A real-world dataset for code generation from webpage designs

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:00.076131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.344168Z digest=sha256:81ab104d5faa4c453b9dda8d4660818e9a6db0800024a89e016f0aab07c185cb

Observation d5b66df0-df1a-4725-8add-002146646b19 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Measuring Massive Multitask Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.353378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.353378Z digest=sha256:6981dff6032e633ea30a38c5fda7931ce02406da09d74d504180b787646e1746

Observation 10c698a1-371a-4203-831b-a0d9b6bb1e55 · outbound

This paper cites TrueSkill : A Bayesian skill rating system.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools TrueSkill : A Bayesian skill rating system

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:00.037332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.362611Z digest=sha256:e539924ca73b5268480c08303782eaf34162cfc32d800f87cc2029c5d9556596

Observation 805bf1fd-3454-406d-82a7-831e5ee645b0 · outbound

This paper cites CLIPScore : A reference-free evaluation metric for image captioning.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools CLIPScore : A reference-free evaluation metric for image captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.369473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.369473Z digest=sha256:d9bad8513202d41dca274c97f986c54135a0dab34b59979dc6b7ea09ba15a34d

Observation d096ce4d-9172-4431-ad5a-b9ce72a23dc5 · outbound

This paper cites GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.384528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.384528Z digest=sha256:4619b60cd85bad68f50803d6979633e988d40b49724be359c064eddd80536314

Observation b33c70c1-dab6-4d35-ad5d-cc14f7848e0a · outbound

This paper cites Jankowski, J.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Jankowski, J

Reference 17

Resolution
verified exact
doi, observed 2026-08-05T15:09:58.131424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.399939Z digest=sha256:caff78bde01f074ba3d2aa3a8e390399e0a1d4fe4b3efc5a7084c5d86d055d86

Observation 1af8fce2-5fe6-41c9-8579-99bde460859d · outbound

This paper cites Jayasumana, X.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Jayasumana, X

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:00.007654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.418581Z digest=sha256:531a90b02bf291487bd8b308bf5f945adec595a99fec786c5b6abff35aac581e

Observation d6959b45-c4ac-4162-91f4-e7a3ecb6b2bf · outbound

This paper cites Kirstain, A.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Kirstain, A

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.976421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.428490Z digest=sha256:27a5e3cc3a610d5ceb80ac3f89ef9922393608171799cd87ce6e7bcdd5b65b2a

Observation 11f51ac6-fab9-401d-9885-932ddb207ef5 · outbound

This paper cites Assessing dimensions of perceived visual aesthetics of web sites.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Assessing dimensions of perceived visual aesthetics of web sites

Reference 20

Resolution
verified exact
doi, observed 2026-08-05T15:09:58.093488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.435107Z digest=sha256:f31f990374790adf7602d2681be0f847ac3dc0b7c3f7721f51156bacea051740

Observation 4b46df08-7fed-4de0-8419-9a51b9f944c7 · outbound

This paper cites Holistic Evaluation of Text-To-Image Models.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Holistic Evaluation of Text-To-Image Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.445683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.445683Z digest=sha256:9601ccf80c557a5b142370f8f729fbe77d48ff4b1320ad372e0421a7d05440ae

Observation d375a1df-5c6f-4bfc-b978-cfe80cc22b82 · outbound

This paper cites Lu, Zhe Lin, Hailin Jin, Jianchao Yang, and James Z.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Lu, Zhe Lin, Hailin Jin, Jianchao Yang, and James Z

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.454500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.454500Z digest=sha256:18a21a0571c0c5ca523a9a15e80b7e5640c99aad612e355789e10d8d900e8bef

Observation cf20d954-4553-47c3-be97-a12dac46b1e1 · outbound

This paper cites WebGen-Bench : Evaluating LLMs on generating interactive and functional websites from scratch, 2025.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools WebGen-Bench : Evaluating LLMs on generating interactive and functional websites from scratch, 2025

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.956095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.473086Z digest=sha256:a651db7b640fe2ffe353359c1b95e6e77b72079eace1bfcc4ed44033c3b64315

Observation ef14ae84-7beb-4282-ba38-330ec87c57ff · outbound

This paper cites FinanceQA : A benchmark for evaluating financial analysis capabilities of large language models, 2025.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools FinanceQA : A benchmark for evaluating financial analysis capabilities of large language models, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.581827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.483190Z digest=sha256:5e0c428b6aa22a10ffb20717b7237b00a0772e01fba2d0148b8849ee71c78e32

Observation 1aa2f16e-85b4-4548-b80f-cd5788622e03 · outbound

This paper cites Moshagen and M.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Moshagen and M

Reference 25

Resolution
verified exact
doi, observed 2026-08-05T15:09:58.052019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.493486Z digest=sha256:da217f84395fc83e03e2ec741159f331f806959e730f469988b5fc21a1b74b84

Observation 5a1b4542-2893-402d-9ec4-9bc09ed90dfe · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 26

Resolution
verified exact
doi, observed 2026-08-05T15:09:58.001466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.498765Z digest=sha256:a0b74a9023740e378ca581e4cf976d5de83a7c7ea91e2ba96a4a7cef4222f8fb

Observation a1a29273-0c60-4209-9cab-56af46a901c0 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 27

Resolution
verified exact
doi, observed 2026-08-05T15:09:57.968857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.508823Z digest=sha256:b761f6b59ebd85fb71763802a892053ddbb69ef9a9ca642904e4ffa2857ec785

Observation 7a3658b6-317a-43be-be2b-509af2ea6a0a · outbound

This paper cites Color compatibility from large datasets.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Color compatibility from large datasets

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.521078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.521078Z digest=sha256:ee0d9083fd12ea8b970a6f93795b91135efbed392a72bf6e310272ca5899fa7f

Observation 3930f5c6-49c0-4fe0-8bda-23a31511888b · outbound

This paper cites Comparative judgement for assessment.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Comparative judgement for assessment

Reference 29

Resolution
verified exact
doi, observed 2026-08-05T15:09:57.944022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.544692Z digest=sha256:97bd4a6aa60b1f88ad70efcd737df2efbe02b450fae0703d0e01368dd897aa0d

Observation 4a059356-abf0-4b3f-8645-fb68b21e424d · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Learning Transferable Visual Models From Natural Language Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.551803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.551803Z digest=sha256:4aea94c89e081c69f1966a586460d8479ce23a9e47c99c1f58cab932b156fe14

Observation 780ed652-e905-4c07-afd6-8047d7423360 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.560815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.560815Z digest=sha256:0ae3b605ede13debe3ca37b76ca1c4396466601efa40c689dffd2a29b4f0b4ab

Observation 726bf2ae-d03a-4edd-a1fa-6f783aa1e233 · outbound

This paper cites Robins and J.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Robins and J

Reference 32

Resolution
verified exact
doi, observed 2026-08-05T15:09:57.890926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.567210Z digest=sha256:28d6bc8a3bec00095d8341692f55b15d9aaa7c7f75cce0ea0a17432a8675d18c

Observation b6e395de-eed0-4f80-bef3-404ac0bf52c9 · outbound

This paper cites Improved techniques for training GANs.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Improved techniques for training GANs

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.528856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.578034Z digest=sha256:ce6f9bb7ed9d20be645c9ce70084c280c783617166ff164d4197affd5deeb27f

Observation 00f33793-f036-4738-b5a9-500883cbaa17 · outbound

This paper cites Design2code: Benchmarking multimodal code generation for automated front-end engineering.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Design2code: Benchmarking multimodal code generation for automated front-end engineering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.486372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.587542Z digest=sha256:e5f22ca38dbbcfa69d93a0d78949954e271b3cc3d39000c1461d63e8ce4e3825

Observation 4bc571d7-7ab6-4154-810c-63e8604c644e · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.600054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.600054Z digest=sha256:28c6c55f62226a40000674e14243843fc69c718ff9bb643f2dec7689e8e63c3d

Observation 7974c531-9c2a-4cf1-8f4c-b6ab4c679265 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.616198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.616198Z digest=sha256:bb56b0c5cb5f733d7925cedf799ee318e7bfce7c9f64fc93bd5f46a605096d19

Observation 59b385df-0352-4354-93e7-a1711bde6633 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:09:59.458898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.632179Z digest=sha256:cd23da0e3e035f0c5d0158fd2d8fdef8dd6a8944f7bb288ed51f9b7b28c32d7c

Observation 9e84e18a-f110-41fd-9462-3f84d46f5535 · outbound

This paper cites Whitehouse and A.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Whitehouse and A

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.411804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.660119Z digest=sha256:e5677cf07800472fe08ccce5631fdc3b83a83f6f36d8081283799c2a4c1a6572

Observation 33c07487-797b-403e-b788-78158d38fad2 · outbound

This paper cites The best AI website builders in 2025.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools The best AI website builders in 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.346581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.666058Z digest=sha256:0150f7a9f92735ebc1061ef922619f4666508465a7c18088c42de6b3c82d4500

Observation 4a486708-4a75-4fc0-b6ee-57bd492e3288 · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:09:59.308828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.676210Z digest=sha256:cdfda12f9da6e1a451360eca5eb560990803a85d6e4d32246c0154e57aa8ce43

Observation 9787ea74-1043-441b-b32d-316a22f13fbc · outbound

This paper cites an unresolved cited work.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:09:59.272758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.685514Z digest=sha256:e6b256b2ebd5fcb7fee470f98364eff853da9dc5ce9a221f9e7e144d13b10fdf

Observation 0a914eb3-5578-49f3-9ac9-f97c4e4a3d6c · outbound

This paper cites Xing, Xiaodan Liang, and Zhiqiang Shen.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Xing, Xiaodan Liang, and Zhiqiang Shen

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.218374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.702743Z digest=sha256:bd22fcac87ecc99199e4ab11abb1c2c39e76c429848aed7f3e9727cd971283b9

Observation 11547017-77de-474a-9487-e17c54314058 · outbound

This paper cites HellaSwag : Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4791--4800, 2019.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools HellaSwag : Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4791--4800, 2019

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.714904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.714904Z digest=sha256:2a23f354dd0ca2359d148ffd0650cbc1cc17e47fa82ab69d1d065893940044f0

Observation 1ddf4ede-2689-4c44-aa82-7a89064a3b35 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.731423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.731423Z digest=sha256:e36a2596a501c10f9b4767ea228cc3478b23a3f288012e4d457f2311005e84bf

Observation 060cab31-313c-40c3-9f4c-a3b8396c4f00 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T15:09:57.740219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:09:57.740219Z digest=sha256:10a8762dd3fa6c7b42600cb0eee6cba23590909078aedeb61df5964212a52819

Observation cde96057-7635-47b3-8f39-15b200a6330d · outbound

This paper cites Frontendbench: A benchmark for evaluating LLMs on front-end development via automatic evaluation, 2025.

UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools Frontendbench: A benchmark for evaluating LLMs on front-end development via automatic evaluation, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:09:59.186704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:09:57.752216Z digest=sha256:9054718a768931d4184f8693b1d3878f10a9333a6e951ee615cff050101dcbbe

Pith citing papers

Observation 7f67039e-6fbf-438a-ab82-ba6c45131e25 · inbound

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation cites this paper.

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.572734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:24:06.863653Z digest=sha256:1011552f6f142ab8653e5a8cb0c89a32a9ed98ece48b9abfc052f818a06d8ee9