Pith. sign in

Paper Citation Record · LEDGER

WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2505.03733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03733 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:49:32.361437Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T14:27:04.294250Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0afcb623-f615-4c20-98b2-6e90b041a284 · inbound

Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing cites this paper.

Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:32.361437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:32.361437Z digest=sha256:2e67a9a35479571cfb543c0d57b6ab00a7cd7cfa12e3cd222cdb8adde7a5c234

Observation 13694a49-581d-4ed8-8a05-2ccb34338de0 · inbound

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software cites this paper.

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T05:18:44.114748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:18:44.114748Z digest=sha256:f766cf75cbb9fc02c5ac28499f8358418ce1152c9f103025faa5f534a0826150

Observation 01469e5f-6421-464a-afdc-22348cfa216c · inbound

MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation cites this paper.

MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:40:19.232153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T11:36:04.456245Z digest=sha256:74c40526be448734933ddd6ecc088e9f6e4fac092ea77f1ff9ed5b6fe9670c0e

Observation 9e1457be-082c-483c-a7e2-1045ebf47852 · inbound

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning cites this paper.

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:19:46.729894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:18:18.340391Z digest=sha256:36aaad52b0f99fce5ed9dfe3e37320a663a24c0d3a5471290766cafaf42dd083

Observation 4ff1504d-f614-4b47-8044-781daa8d3c33 · inbound

SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies cites this paper.

SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:08.931833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T16:24:49.710882Z digest=sha256:16c45628d2acfac0682431cfc230cd7e0dfc5851120a1cf85c72e2b3ff375117

Observation 3a3d8263-d702-4510-999f-85fc96de1a2e · inbound

GameGen-Verifier: Parallel Keypoint-Based Verification for LLM-Generated Games via Runtime State Injection cites this paper.

GameGen-Verifier: Parallel Keypoint-Based Verification for LLM-Generated Games via Runtime State Injection WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:55.335095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:13:39.714918Z digest=sha256:53cafb1bab6e4bcad082bfdb672790982e047365c598320e2420d065af385da4

Observation 6a1925d2-145a-4166-ba3c-b26ec5afe4cc · inbound

From Runnable to Shippable: Multi-Agent Test-Driven Development for Generating Full-Stack Web Applications from Requirements cites this paper.

From Runnable to Shippable: Multi-Agent Test-Driven Development for Generating Full-Stack Web Applications from Requirements WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:22:52.285629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T23:22:10.879878Z digest=sha256:e2f4a7fe19e5e7f92936247f974cd0173c0f6190614c31127ef73d22b261c887

Observation 4ea8f743-b231-449e-9281-4049a6cd736d · inbound

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents cites this paper.

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:12:51.593866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T23:09:56.130802Z digest=sha256:f7c1c1c42c29931f3df67a1c74c1e8f33705449215960dbb583bfed15d2f4b8a

Observation 00ca0e8f-e230-4e13-8c8d-e9176e509aa0 · inbound

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents cites this paper.

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:08:17.896885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T13:05:38.485058Z digest=sha256:f06c25d651f6b7095af85a16c9b1a6b2eac37c5609ca1e7b73b515cfad548063

Observation 2aa4923b-c732-4508-ac48-fd53dfb44ba2 · inbound

SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering cites this paper.

SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T22:43:14.965403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T22:43:13.251812Z digest=sha256:742e24f8814258f20c333ad9293d588631bdd6a75fdf9f104da54032b7e696e3

Observation ed6e5465-653c-494b-a89a-067a6d0a7888 · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:28:17.235684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:24:06.062957Z digest=sha256:cf60974bdd9d8d49a3963219aacc2c8d7019772259cc9e93e18e427fe9ff0f66

Observation 1e345fc4-68bd-4168-b85d-95addb9972d9 · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:45:23.245448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:45:04.573722Z digest=sha256:c8e9fc7b10f8f68695cdae366da08f07620d5224034edefeb19e77e914fb5962

Observation 5de4f68f-7581-498e-836f-ae8fa8be7db2 · inbound

HTMLCure: Turning Browser Experience into State Guided Repair for Interactive HTML cites this paper.

HTMLCure: Turning Browser Experience into State Guided Repair for Interactive HTML WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:03:34.059095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T15:56:24.222373Z digest=sha256:e87ce697c9ac52745457ced6b1ffa478489318f7937da9d0b8f6ff50f8ba0463

Observation f7039c47-6d8b-4b1c-9bb6-b6fde9cde118 · inbound

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation cites this paper.

Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.533156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T07:24:06.863653Z digest=sha256:d6f318a2588e508ecd64357b807719b2ca332290c76f9524b2eb9d52627d0f05

Observation 3bcb69ed-e434-4381-bd5c-bc7b0d099c7b · inbound

QASM-Eval: A Dataset to Train and Evaluate LLMs on OpenQASM-3 Beyond Quantum Circuits cites this paper.

QASM-Eval: A Dataset to Train and Evaluate LLMs on OpenQASM-3 Beyond Quantum Circuits WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:35:33.308090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T08:35:24.708864Z digest=sha256:1c5325acdaba905f98447058fb4af138d68b9077b051f3dabf4ad56c350a7cdc

Observation 4484d5b1-5e5c-41fc-bbfb-cb229759fd2c · inbound

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications cites this paper.

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:42:36.199841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T18:53:18.645984Z digest=sha256:bae4b999f8b2463b7b54914f2ec243ef892450d37e49444503f33e447a7c16b3

Observation 6f1ad68c-70e0-49cf-8b10-a98aaeb1b1b7 · inbound

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement cites this paper.

Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:27:04.296226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T00:26:22.041924Z digest=sha256:9a3a93e92858b8300befa8c6a1be1e3da71158c61129704a7db571cf8ca5584b

Observation df41b1aa-0ac2-48d4-9ca5-8afa6e89fc61 · inbound

Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation cites this paper.

Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T05:40:26.231504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:40:26.231504Z digest=sha256:5ee71715b1e8b23125a5ef06ce389422ae749992388a4fe994d4eaa0c634bbd6