Pith. sign in

REVIEW 4 major objections 5 minor 25 references

A language-model agent can satisfy every functional requirement in a task and still deliver an unusable or untrustworthy artifact; this paper argues that evaluation must therefore report risk determinations separately from functional comple

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 03:32 UTC pith:MDDP43SB

load-bearing objection SQBench is a useful, honestly reported benchmark whose key number (113) needs external validation before you trust it. the 4 major comments →

arxiv 2607.23123 v1 pith:MDDP43SB submitted 2026-07-25 cs.AI

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

classification cs.AI
keywords language-model agentsbenchmarktask deliveryconstrained workflowsverifiable deliverablesrisk-aware evaluation10D Risk MatrixStrict Pass
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

SQBench is a benchmark for judging language-model agents on the deliverable they produce inside a constrained workflow, not just on the answer they write. Each of 220 tasks asks an agent to process materials, use tools, and hand in a specified artifact; scoring first checks functional Completion and then applies a 10D Risk Matrix that records independently evidenced failures such as unverifiable citations, format violations, and unauthorized actions. A Strict Pass requires both Completion = 1 and Risk Penalty = 0. Across 27 model configurations, 113 of 2,348 functionally complete results failed strictly because of such risks, and domain-constrained L3 tasks show a shared collapse (mean Strict Pass 18.5%). The paper's central point is that delivery quality has a risk dimension that answer-style metrics miss, and that risk determinations should be reported separately.

Core claim

On SQBench v1.0, the central claim is that functional completion and reliable delivery are different quantities. The benchmark treats the verifiable deliverable as the unit of evaluation and records two independent signals: Completion, which captures whether the required files, fields, formats, and behaviors are present, and Risk Penalty, which is computed from predefined triggers in a 10D Risk Matrix that require independent evidence in the artifact or trace. A task is Strict Pass only when Completion equals 1 and Risk Penalty equals 0. Empirically, 4.8% of results that fully satisfied functional requirements (113 of 2,348) triggered nonzero risk penalties, mostly on L2 composite workflows,

What carries the argument

The 10D Risk Matrix: a fixed set of ten delivery-risk dimensions (factual/provenance hallucination, format/instruction violation, compliance/safety overreach, inefficiency/resource misuse, tool-state hallucination, premature termination/goal forgetting, epistemic boundary/escalation, sycophancy/deception, process opacity, context contamination) each with a predefined penalty. Its role is to convert risk into a numeric penalty derived only from independent evidence in the artifact or trace, so that Performance = max(0, Completion − Risk Penalty) and Strict Pass requires completion with zero penalty.

Load-bearing premise

The load-bearing premise is that the 10D risk-trigger rules in Section 3.6 correctly identify genuine delivery failures from independent evidence; if those triggers are mis-specified, the 113 'complete but not Strict Pass' cases that support the paper's main claim could be artifacts of the rubric rather than real problems.

What would settle it

Take the 113 results with Completion = 1 and Risk Penalty > 0 and have independent human experts score each deliverable for real-world usability without seeing the risk flags; if experts judge most of those deliverables as actually usable (e.g., the unverifiable citations are verified, the resource use is within normal range, or the format violation is easily fixed), then the risk dimension is not separating completion from delivery quality, and the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Reporting functional Completion alone understates failure: within SQBench v1.0, 4.8% of fully completed tasks carry a risk penalty that blocks a passing grade.
  • Risk-aware evaluation changes how model configurations compare; the gap between mean Completion and mean Performance across 27 configurations averages 7.5 percentage points.
  • Domain-constrained delivery (L3) is a collective weakness: all 27 configurations score lower on L3 than on L1 and L2, with a mean Strict Pass of 18.5%.
  • The decomposition into Completion, risk evidence, and Strict Pass is auditable at task level, since each triggered dimension leaves a trace or artifact that can be inspected.
  • Because every v1.0 penalty is positive, any triggered risk dimension excludes a Strict Pass, making the binary distinction between function and risk easy to reproduce and audit.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the pattern generalizes, the 4.8% share could be larger in higher-stakes workflows where artifacts are consumed by downstream processes, because format and provenance failures matter more there than in research-style tasks.
  • The 10D penalty weights are tunable; calibrating them against external incident severity or downstream usability could turn the binary Strict Pass into a more graded measure of delivery trustworthiness.
  • The same two-signal design (functional completion plus separate risk evidence) could be layered onto existing agent benchmarks without rebuilding their task sets, by adding an audit pass over their outputs.
  • Re-running the 113 complete-but-risky cases with human experts as judges would test whether the rubric's risk triggers match actual delivery failure rather than rubric artifacts.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces SQBench, a 220-task benchmark for evaluating language-model agents on production-oriented task delivery. Each task requires the agent to process input assets, use tools, and produce a specified deliverable in a controlled environment. The evaluation computes functional Completion, then applies a 10D Risk Matrix with predefined penalties to derive Risk Penalty, Performance, and a binary Strict Pass. The authors evaluate 27 model configurations with one run per configuration-task pair and report that 113 of 2,348 functionally complete results (4.8%) fail Strict Pass because of recorded risks, primarily D4 (Inefficiency and resource misuse). The paper interprets this as evidence that functional completion alone does not fully characterize delivery quality and that risk determinations should be reported separately. The manuscript is careful about scope and openly lists important limitations, including unvalidated risk dimensions, a single LLM judge, hidden tasks, and single-run observations.

Significance. If the risk layer is valid, SQBench makes a useful contribution by shifting evaluation from answer correctness or environment state to artifact-level delivery quality, with an explicit separation of functional Completion from risk-adjusted Strict Pass. The paper has notable strengths: the metrics are prespecified, the risk penalties are explicit rather than fitted, a fixed public package with aggregate results and verification code is provided, and the weight-sensitivity analysis in Appendix E.2 shows that the top configuration is stable under alternative aggregations. The L3 finding—that every configuration performs worse under domain constraints—is a straightforward descriptive result within the task set. However, the central empirical claim that functional completion undercharacterizes delivery quality rests on the 113 'complete but not strict' cases, and the current manuscript does not establish that these cases correspond to genuine delivery failures rather than rubric artifacts. The missing validation of risk triggers and the LLM judge, together with the hidden task set, are the main barriers to accepting the central claim. With validation, the contribution would be significa

major comments (4)
  1. [§3.6/Table 1; §5.3/Table 3] The central claim rests on the 113 complete-but-not-strict results, 92 of which trigger D4. D4 and D9 are process-level judgments ('repeated failed attempts,' 'ineffective loops,' 'process opacity') that lack external calibration. Section 6 concedes that the dimensions and penalties are 'versioned design choices' with no empirical calibration, no judge agreement, and no human-expert baseline. As defined, a run with a correct deliverable but a visible retry loop can receive a nonzero penalty (D4 0.3; D9 0.2), so the 113 cases may be artifacts of the rubric. The paper needs validation—human-expert annotation of risk triggers on a sample, agreement metrics, and an error analysis showing that the 113 cases are genuinely unusable or unauditable. Without this, the main empirical claim is unsupported.
  2. [§3.5/§4.2/§6] 107 of 220 tasks use a single LLM judge (modelstudio/qwen3.5-122b-a10b) for semantic requirements, and the paper reports no inter-judge agreement, task-designer agreement, or human baseline. Because Completion=1 is a necessary condition for the 113 'complete but not strict' cases, the split between Completion and Strict Pass is sensitive to judge leniency or strictness. The authors should report agreement statistics and, ideally, recompute the 113-case result under alternative judges or on a human-scored subset. This is load-bearing for the paper's central claim, not just a quality-of-scoring concern.
  3. [§3.3/§6/Data Availability] The hidden-task design prevents external verification: official prompts, input assets, reference answers, scoring scripts, the model-by-task matrix, and raw traces are withheld. The public package includes only aggregate statistics and three synthetic examples. Given that the benchmark author's organization also develops and operates the leaderboard, the lack of independent audit access is a governance risk for a benchmark that claims auditability as a design goal. The paper should provide a concrete review mechanism—e.g., controlled evaluation access or release of anonymized task-level scores and traces for the 113 disputed cases—so that external researchers can verify the central claim.
  4. [§4.1/§6] One run per configuration-task pair produces point estimates with no uncertainty. The headline numbers—4.8% complete-but-not-strict, 60.5% top Weighted Pass@1, and the L3 mean of 18.5%—are single observations. Repeated runs, random seeds, or service-version changes could shift the 113 count, and the paper's own limitations section acknowledges this. Since the 113-case difference is a small fraction of total results, the authors should either report bootstrap or repeated-run variability on the risk-triggering subset or explicitly frame the 113-case count as a single-run observation that does not yet establish the central claim at the population level. This would substantially strengthen the paper.
minor comments (5)
  1. [§3.7] The equivalence Strict Pass ⇔ Completion = 1 and Risk Penalty = 0 is stated unconditionally, but it depends on all v1.0 penalties being positive. Please add the qualifier so readers understand that the binary equivalence is a v1.0 design property.
  2. [§5.3] The trigger counts among the 113 complete-but-not-strict results sum to exactly 113 (D4 92 + D10 8 + D2 6 + D1 5 + D7 2). Since the paper says one result may trigger multiple dimensions, please clarify whether these are unique-result counts or a case where no overlapping triggers occurred; otherwise the arithmetic looks inconsistent.
  3. [Figure 4] The caption refers to the 'gap' between Completion and Performance but does not define it. Please state the gap as mean Completion minus mean Performance across all tasks for each configuration.
  4. [§5.1] The statement that 'even the highest-scoring configuration fails to satisfy both Completion = 1 and Risk Penalty = 0 on nearly half of the tasks' would benefit from citing the corresponding Simple Pass@1 (54.5%) directly in the sentence, since the connection to 'nearly half' is otherwise implicit.
  5. [Appendix D.1] The judge model is fixed to modelstudio/qwen3.5-122b-a10b; because judge choice is known to affect evaluation, consider reporting the judge model as part of the benchmark configuration and adding a sensitivity check with at least one other judge.

Circularity Check

0 steps flagged

No significant circularity: Strict Pass is explicitly defined from Completion and Risk Penalty; the 113-case gap is an observed count, and the key limitation (unvalidated 10D triggers) is an external-validity concern, not a circular derivation.

full rationale

The derivation chain in SQBench is short, explicit, and does not reduce to its inputs. Completion is measured from automated checks and rubric-based judge scores; Risk Penalty is the sum of penalties for triggered 10D dimensions; Performance = max(0, Completion - Risk Penalty); and Strict Pass is defined as Completion = 1 and Risk Penalty = 0 (Section 3.7, Appendix D.1). No parameter is fitted to the headline result, and no prediction is renamed from a fit. The 113/2,348 cases with Completion = 1 that fail Strict Pass are empirical counts from the task runs: if no risk trigger ever fired, that number would be zero, so the result is not forced by the equations alone. The paper is transparent that the risk dimensions and penalties are 'versioned design choices' and are 'not empirically calibrated against external incidence, relative severity, or optimal penalty magnitude' (Section 6). That is a validity limitation about whether the 10D triggers identify genuine delivery failures, not a circularity in the internal derivation. The references are external benchmark and method citations; there is no load-bearing self-citation or imported uniqueness theorem. The author's organization maintaining the hidden benchmark and leaderboard is a governance and audit concern, not an internal circularity. Accordingly, no circular step is identified.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 1 invented entities

The benchmark's conclusions depend on the validity of its risk taxonomy, the reliability of LLM-judge scoring, and the representativeness of a hidden 220-task set. The paper explicitly scopes its claims to v1.0 and discloses the lack of calibration, so these ledger entries are assumptions rather than hidden flaws, but they are unverified.

free parameters (3)
  • 10D risk penalty magnitudes = D1=0.8, D2=0.4, D3=1.0, D4=0.3, D5=0.5, D6=0.6, D7=0.6, D8=1.0, D9=0.2, D10=0.5
    Hand-set penalty values in Table 1. They determine the reported Performance gap (4.2-11.3 points) but not the binary Strict Pass outcome, since all penalties are positive.
  • Weighted Pass@1 layer weights = 0.2 / 0.6 / 0.2 for L1/L2/L3
    Prespecified aggregation in Section 3.7. Affects ranking; sensitivity check shows top configuration stable, middle ranks shift.
  • Task-specific Completion weights for judge/automated scores = task-specific, not disclosed per task
    Equation in D.1: Completion = (w_A A + w_J J)/(w_A + w_J). The weights are fixed in advance but the values are not published for each task, making Completion not independently reproducible.
axioms (4)
  • domain assumption The 10D Risk Matrix dimensions and trigger conditions are a valid characterization of delivery risk.
    Section 6: dimensions and penalties are versioned design choices not calibrated against external incidence or severity.
  • domain assumption The LLM judge with predefined rubrics produces valid semantic quality and risk-evidence judgments.
    Section 6: no judge agreement or human-expert baseline; cited biases in [24,25] acknowledged.
  • domain assumption A single run per configuration-task pair is sufficient for the descriptive claims.
    Section 4.1 and 6: no repeated sampling or uncertainty estimation; paper limits claims to this task set.
  • ad hoc to paper The hidden task set is representative enough of production-oriented workflows to support the stated conclusions.
    Section 6: tasks are abstractions, not a statistically representative sample; conclusions explicitly scoped to v1.0 tasks.
invented entities (1)
  • 10D Risk Matrix dimensions (D1-D10) no independent evidence
    purpose: Classify delivery risks and trigger Risk Penalty when independently evidenced.
    New constructs introduced by this paper; no external validation, calibration, or inter-rater reliability data provided.

pith-pipeline@v1.3.0-alltime-deepseek · 11774 in / 11985 out tokens · 107530 ms · 2026-08-01T03:32:23.368652+00:00 · methodology

0 comments
read the original abstract

Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced within a constrained workflow as the unit of evaluation. We introduce SQBench, a benchmark for evaluating production-oriented task delivery by language-model agents. SQBench v1.0 contains 220 standardized tasks organized into L1 atomic capabilities, L2 composite skills, and L3 business scenarios. Each task requires an agent to process input assets, use available tools, and produce an explicitly specified deliverable. The evaluation first computes functional Completion and then derives Risk Penalty and Performance from independently evidenced triggers in a 10D Risk Matrix. A Strict Pass requires Completion = 1 and Risk Penalty = 0. We evaluate 27 model configurations under a common protocol, with one run per configuration-task pair. The highest prespecified Weighted Pass@1 is 60.5%. Mean Strict Pass@1 on L3 is 18.5%, and every configuration performs worse on L3 than on both L1 and L2, indicating that delivery under domain constraints is a shared weakness within the current task set. Of 2,348 results with Completion = 1, 113 (4.8%) fail the Strict Pass criterion because of risks such as unverifiable citations, inappropriate resource use, or format violations. These results show that functional completion alone does not fully characterize delivery quality and that risk determinations should be reported separately.

Figures

Figures reproduced from arXiv: 2607.23123 by Summer Sun (Shaqiu Community).

Figure 1
Figure 1. Figure 1: SQBench task execution and evaluation workflow. Each task runs under a fixed task version, initial assets, tool permissions, resource limits, and scoring rules. The execution trace and final deliverable jointly support scoring and error analysis. 3.1 Design principles SQBench task design and evaluation are guided by four principles: workflow abstraction, deliverable verification, risk assessment, and withi… view at source ↗
Figure 2
Figure 2. Figure 2: Three-layer task structure of SQBench. From bottom to top, the layers represent foundational execution, composite workflows, and domain-constrained delivery. They describe capability scope and contextual constraints rather than a difficulty scale. 3.5 Dataset composition SQBench v1.0 contains 220 standardized tasks: 100 L1 tasks, 60 L2 tasks, and 60 L3 tasks. L3 spans finance, the industrial sector, health… view at source ↗
Figure 3
Figure 3. Figure 3: Distribution of layerwise Strict Pass rates. Blue circles, cyan squares, and orange triangles represent L1, L2, and L3 Pass@1. Thin lines connect results from the same configuration, and black diamonds show layer means. L3 is a shared weakness. Mean Strict Pass@1 is 39.8% on L1, 53.1% on L2, and 18.5% on L3, and every configuration performs worse on L3 than on both L1 and L2. Claude Opus 4.8 (Max) has the … view at source ↗
Figure 4
Figure 4. Figure 4: Model scores before and after risk adjustment. Open points show mean Completion and filled points show mean Performance after Risk Penalty. Connecting segments represent the reduction; configurations are sorted by gap. 5.2 L3 results by industry and task type [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Strict Pass rates by L3 industry and task type. Each cell is the mean Strict Pass rate of 27 configurations on the corresponding five-task subset. The rightmost column and bottom row show industry and task-type means. By task type, quantitative audit has a mean Strict Pass rate of 10.7%, below complex business at 20.2% and compliance and risk control at 24.6%. Twenty-five of the 27 configurations perform b… view at source ↗
Figure 6
Figure 6. Figure 6: Layerwise trigger rates for the 10D Risk Matrix. Each cell shows the fraction of model-task results in a layer that trigger the corresponding dimension. A result may trigger multiple dimensions. D4 has the highest total trigger count across the three layers; [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: presents an observed result that is functionally complete but does not meet Strict Pass. DeepSeek￾V4-Pro (High) creates the required file, uses the required tools, satisfies the report structure, and obtains Completion = 1.00 on an L2 market-research task. Risk assessment finds unverifiable or apparently fabricated citation links, triggering D1 and a penalty of 0.80. Performance is therefore 0.20 [PITH_FU… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

25 extracted references · 15 linked inside Pith

  1. [1]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. InInternational Conference on Learning Representations, 2021

  2. [2]

    Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.arXiv preprint arXiv:2206.04615, 2022

    Aarohi Srivastava et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.arXiv preprint arXiv:2206.04615, 2022

  3. [3]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level google-proof Q&A benchmark.arXiv preprint arXiv:2311.12022, 2023

  4. [4]

    GAIA: A benchmark for general AI assistants

    Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. GAIA: A benchmark for general AI assistants. InInternational Conference on Learning Representations, 2024

  5. [5]

    AgentBench: Evaluating LLMs as agents.arXiv preprint arXiv:2308.03688, 2023

    Xiao Liu et al. AgentBench: Evaluating LLMs as agents.arXiv preprint arXiv:2308.03688, 2023

  6. [6]

    WebArena: A realistic web environment for building autonomous agents.arXiv preprint arXiv:2307.13854, 2023

    Shuyan Zhou et al. WebArena: A realistic web environment for building autonomous agents.arXiv preprint arXiv:2307.13854, 2023

  7. [7]

    OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments

    Tianbao Xie et al. OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. InAdvances in Neural Information Processing Systems, 2024

  8. [8]

    MLAgentBench: Evaluating language agents on machine learning experimentation.arXiv preprint arXiv:2310.03302, 2023

    Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec. MLAgentBench: Evaluating language agents on machine learning experimentation.arXiv preprint arXiv:2310.03302, 2023

  9. [9]

    ToolQA: A dataset for LLM question answering with external tools

    Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang. ToolQA: A dataset for LLM question answering with external tools. InAdvances in Neural Information Processing Systems, 2023

  10. [10]

    ToolLLM: Facilitating large language models to master 16,000+ real-world APIs.arXiv preprint arXiv:2307.16789, 2023

    Yujia Qin et al. ToolLLM: Facilitating large language models to master 16,000+ real-world APIs.arXiv preprint arXiv:2307.16789, 2023

  11. [11]

    Toolformer: Language models can teach themselves to use tools

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems, 2023

  12. [12]

    API-Bank: A comprehensive benchmark for tool-augmented LLMs.arXiv preprint arXiv:2304.08244, 2023

    Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. API-Bank: A comprehensive benchmark for tool-augmented LLMs.arXiv preprint arXiv:2304.08244, 2023. 16 SQBench Summer Sun

  13. [13]

    Jimenez et al

    Carlos E. Jimenez et al. SWE-bench: Can language models resolve real-world GitHub issues? In International Conference on Learning Representations, 2024

  14. [14]

    WorkArena: How capable are web agents at solving common knowledge work tasks?arXiv preprint arXiv:2403.07718, 2024

    Alexandre Drouin et al. WorkArena: How capable are web agents at solving common knowledge work tasks?arXiv preprint arXiv:2403.07718, 2024

  15. [15]

    Xu et al

    Frank F. Xu et al. TheAgentCompany: Benchmarking LLM agents on consequential real world tasks. arXiv preprint arXiv:2412.14161, 2024

  16. [16]

    APEX-Agents.arXiv preprint arXiv:2601.14242, 2026

    Bertie Vidgen et al. APEX-Agents.arXiv preprint arXiv:2601.14242, 2026

  17. [17]

    GDPval: Evaluating AI model performance on real-world economically valuable tasks.arXiv preprint arXiv:2510.04374, 2025

    Tejal Patwardhan et al. GDPval: Evaluating AI model performance on real-world economically valuable tasks.arXiv preprint arXiv:2510.04374, 2025

  18. [18]

    Holistic evaluation of language models.arXiv preprint arXiv:2211.09110, 2022

    Percy Liang et al. Holistic evaluation of language models.arXiv preprint arXiv:2211.09110, 2022

  19. [19]

    Mind2Web: Towards a generalist agent for the web

    Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2Web: Towards a generalist agent for the web. InAdvances in Neural Information Processing Systems, 2023

  20. [20]

    AgentBoard: An analytical evaluation board of multi-turn LLM agents.arXiv preprint arXiv:2401.13178, 2024

    Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He. AgentBoard: An analytical evaluation board of multi-turn LLM agents.arXiv preprint arXiv:2401.13178, 2024

  21. [21]

    Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan.τ-bench: A benchmark for tool- agent-user interaction in real-world domains.arXiv preprint arXiv:2406.12045, 2024

  22. [22]

    As- sistantBench: Can web agents solve realistic and time-consuming tasks?arXiv preprint arXiv:2407.15711, 2024

    Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin, Ofir Press, and Jonathan Berant. As- sistantBench: Can web agents solve realistic and time-consuming tasks?arXiv preprint arXiv:2407.15711, 2024

  23. [23]

    ReAct: Synergizing reasoning and acting in language models

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. InInternational Conference on Learning Representations, 2023

  24. [24]

    G-Eval: NLG evaluation using GPT-4 with better human alignment.arXiv preprint arXiv:2303.16634, 2023

    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-Eval: NLG evaluation using GPT-4 with better human alignment.arXiv preprint arXiv:2303.16634, 2023

  25. [25]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM- as-a-judge with MT-Bench and Chatbot Arena. InAdvances in Neural Information Processing Systems, 2023. 17