{"id":"84c70d68-b91b-4ec2-a8e1-17d0a23a8ab1","arxiv_id":"2501.13120","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"English prompts produce better LLM-designed reward functions for RMABs than Hindi, Tamil, or Tulu prompts, and complex prompts worsen performance and fairness across languages.","lead":"Researchers prompted an LLM-based reward function generator for restless bandits in English, Hindi, Tamil, and Tulu. They found English prompts yield better task performance and fewer unintended fairness gaps, while low-resource languages and complex prompts perform worse.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The English-vs-other-languages comparison rests on unverified prompt translations; if the Hindi/Tamil/Tulu prompts differ in meaning from the English originals, the observed gaps are artifacts, not language effects.","rationale":"The reader's weakest-assumption analysis correctly identifies the translation issue. I agree with that choice over the statistical caveat, because the translation confound threatens the experiment's construct validity: the independent variable is supposed to be language, not prompt content. Without evidence of translation fidelity, the headline result could be an artifact of mistranslation. The paper does include some independent support: a synthetic environment with explicit structural equations and a well-defined fairness metric (DP variance) that makes the experiments reproducible in principle. However, reproducibility is limited by the absence of code, data, and the translated prompts themselves. The statistical weakness (overlapping error bars in Table 3) reinforces the need for caution, but it is secondary to the translation concern. I therefore recommend keeping the verdict conditional: the authors should release verified translations and either provide significance tests or soften the 'significantly better' language.","tokens_in":10322,"tokens_out":4975,"duration_ms":53363,"concrete_test":"Contact the authors for the full set of translated prompts (8 prompts × 3 languages) and have independent native speakers back-translate them into English. Compare each back-translation against the original English prompt for semantic equivalence, specifically checking feature names, quantifiers ('slightly'), and demographic labels. If any translation alters a feature, condition, or strength, mark that prompt as confounded and rerun the experiments with corrected, verified translations. If the English advantage persists under verified translations, the concern is resolved; if not, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that English prompts outperform Hindi/Tamil/Tulu and that low-resource languages create unfairness—presupposes that the non-English prompts are semantically equivalent translations of the English originals. This is unverified. Section 2.2 (Table 2) lists only the English prompts; the actual translated prompts, back-translations, native-speaker checks, and translation-quality metrics are absent. The paper does not state who translated the prompts, with what instructions, or whether translations were validated. For complex prompts like Prompt 6 ('older women with lower incomes who are Kannadiga'), a translation that renders 'Kannadiga' as 'Kannada speaker' or drops 'slightly' could change the intended allocation target. If translations differ in meaning or specificity, the observed performance and fairness gaps are confounded with translation quality and cannot be attributed to language alone. Additionally, Table 3's overlapping standard errors mean the 'significantly better' claim is not statistically supported, but the translation confound is the more fundamental threat to construct validity: even perfect statistics cannot salvage a comparison that is not actually varying language.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how the language of goal prompts affects the DLM algorithm, which uses an LLM (Gemini 1.0 Pro) to design reward functions for restless multi-armed bandits in a public-health-inspired resource allocation setting. The authors construct a synthetic environment with six features, run eight prompts (six increasing in complexity and two rephrased versions) in English, Hindi, Tamil, and Tulu, and evaluate three outcomes: the rate of 'acceptable' reward functions, task performance, and fairness measured by demographic-parity variance. The paper reports that English prompts produce better reward functions and task performance, that phrasing matters, that performance degrades with prompt complexity but less so for English, and that low-resource languages and complex prompts are more likely to induce unfairness along unintended features.","tokens_in":10561,"tokens_out":4419,"duration_ms":53369,"significance":"If fully supported, this would be a valuable early empirical result on language-dependent behavior of LLM-based reward design, with direct implications for equitable deployment of automated resource allocation in multilingual, low-resource settings. The paper addresses an understudied question—non-English prompts and fairness rather than performance alone—and includes a low-resource language (Tulu), multiple prompt complexities, and two correlation strengths. Its strengths include the explicit study of prompt phrasing and the use of a fairness metric beyond task success. However, the central comparisons are not yet supported by the evidence as presented: the 'significantly better' claim lacks inferential tests, the multilingual manipulation is not validated as translation-equivalent, and the task-success/fairness thresholds are not operationalized. These issues are fixable in revision and do not invalidate the research direction, but they block acceptance in the current form.","major_comments":[{"comment":"The claim that LLM-proposed reward functions are 'significantly better when prompted in English' is not supported by any significance test. In Table 3, the reported standard errors overlap for most language pairs: for Prompt 1, English is 0.65 ± 0.15 versus Hindi 0.60 ± 0.15; for Prompt 5, all four languages report 0.10. The paper provides no paired test across the 20 runs, no confidence intervals for differences, and no effect-size measure. Because this claim is the paper's headline result, the authors must report appropriate inferential statistics (for example, paired bootstrap or McNemar-style tests on the per-run acceptability outcomes) or soften the language to descriptive trends.","section":"Abstract and Section 3.1, Table 3"},{"comment":"The English-versus-other-languages comparison presupposes that the Hindi, Tamil, and Tulu prompts are semantically equivalent translations of the English originals, but the paper gives no evidence for this. Table 2 lists only English prompts; the translated prompts, the translation procedure, translator qualifications, back-translations, and any native-speaker validation are absent. Without this information, observed performance and fairness gaps are confounded with translation quality and cannot be attributed to language per se. The authors should include the full translated prompt set and a validation protocol, and ideally a control condition that perturbs the English prompts to quantify sensitivity to wording.","section":"Section 2.2, Table 2"},{"comment":"The 'task success rate' used to support Result 4 is not operationally defined. The text says success occurs when 'the allocations have deviated from the relevant features to the point where the allocation is higher than what it would have been had the allocation been uniform across feature values,' but no threshold, comparison procedure, or code is provided. This makes the complexity-degradation and English-robustness claims non-reproducible. In addition, the authors' own Note in Section 4.3 states that the very poor performance for Prompts 2 and 3 is due to failure on ordered categorical features; because Prompt 2 (two features) is not systematically easier than Prompt 5 (two features plus an inference step) in Table 3, the monotonic complexity story is not supported by the raw acceptability rates. The analysis should separate feature-type effects from complexity effects and precisely define the success criterion.","section":"Section 4.3 and Figure 6/7"},{"comment":"The fairness results depend on two unspecified choices. First, the 'certain threshold' used for the absolute DP-variance counts in Figures 9 and 10 is not reported, so the reader cannot tell whether the plotted differences are robust or an artifact of one threshold value. Second, the relative count (unintended-feature DP variance greater than intended-feature DP variance) may be trivially sensitive to whether the intended feature has high variance by construction. Furthermore, Result 5 is reported as strong for α = 0.2 but 'not as pronounced' for α = 0.8; the paper explains this by correlation propagation but does not provide a sensitivity analysis. The fairness claim 'low-resource languages create unfairness' is therefore conditional on one synthetic correlation structure and should be qualified accordingly, with thresholds reported and confidence intervals or error bars given for the counts.","section":"Section 5, Figures 8–11"}],"minor_comments":[{"comment":"Equation (3) defines P(Y=1) as the unweighted average of group-level allocation probabilities, not the overall allocation rate across arms; this should be clarified because the demographic-parity variance in Equation (2) then compares group probabilities against an unweighted mean rather than the population prevalence.","section":"Section 2.3, Equation (3)"},{"comment":"The caption calls the entries 'acceptable prompt rates,' but the rate is of acceptable reward functions, not prompts; please reword for clarity.","section":"Section 3, Table 3"},{"comment":"The full set of translated prompts is not included in the paper or an appendix; providing them is essential for reproducibility and for any verification of translation equivalence.","section":"Section 2.2"},{"comment":"The feature weight vector is described as 'inspired by' prior work, but the choice of signs and magnitudes is one of several plausible specifications; a sensitivity check over the weight vector would strengthen the claim that the results are not artifacts of the synthetic environment.","section":"Section 2.1, Table 1"},{"comment":"The x-axis label 'increasing complexity' is not numerically defined; consider reporting the number of intended features and the reasoning step for each prompt directly in the figure.","section":"Section 4.3, Figures 6–7"}],"recommendation":"major_revision","confidential_remarks":"The paper would be much easier to evaluate if the authors released the translated prompts, the exact success and fairness thresholds, and the code for the synthetic environment and DLM evaluation. The overlap in author teams with the original DLM work is not itself a problem, but the evaluation would benefit from an explicit statement of what components are reused unchanged versus newly implemented for this study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Key thing to know: the paper asks a good question and reports a plausible cautionary result, but the central comparison is under-powered and confounded. The abstract's claim that English-prompted reward functions are 'significantly better' rests on Table 3, where most pairs overlap by well over a standard error and no significance tests are reported. So that specific word is unsupported. That said, the underlying pattern—English more robust, low-resource languages and more complex prompts more likely to introduce unfairness—is consistent and worth taking seriously.\n\nWhat's actually new: as far as the paper's literature review goes, it is the first to look at non-English prompting for LLM-designed reward functions in RMABs, and it brings fairness into the evaluation. That is a legitimate new application with real practical relevance for public-health deployment in India. The result that phrasing alone changes allocations (Prompts 1 vs 7, 2 vs 8) is a useful, concrete finding. The authors are also honest about the weak performance on ordered categorical features (their note in Section 4.3), which I read as good faith.\n\nWhere it gets soft, in rough order of severity. (1) Translation fidelity is unverified. The paper lists only English prompts and gives no back-translations, native-speaker checks, or translation-quality metrics. The stress-test note is right: if the Hindi/Tamil/Tulu prompts differ semantically from the English originals—say, if 'Kannadiga' is rendered as 'Kannada speaker' or 'slightly' is dropped—then the observed gaps are artifacts of translation, not language. This is the most serious threat to construct validity. (2) The 'significantly better' phrasing in the abstract is not supported by the reported statistics. With overlapping error bars, the honest claim is 'appears better in these runs.' (3) The synthetic environment uses hand-set structural equations and feature weights. That is acceptable for a stress test, but it limits generalization. (4) No code or data is released, so the 20-run results cannot be checked. That is fixable but real.\n\nThe citation pattern looks fine. The authors build directly on Behari et al. and Verma et al., and shared authorship there is not a problem; DLM is the natural baseline and they say so. Related work is adequate.\n\nWho benefits: anyone working on LLM-designed rewards for sequential decision-making, and anyone deploying such systems in multilingual, low-resource settings. The paper is not a definitive demonstration, but it is a useful alarm bell. It deserves a serious referee: the question is socially important, the experimental outline is reasonable, and the flaws are fixable. I would send it to review with a request for major revisions—add significance tests or reword the claims, release translated prompts and code, and verify translation quality with native speakers or back-translation.","headline":"The paper's headline claim of significantly better English performance is not backed by its own statistics, but the question is real, the study is honest, and it deserves a rigorous peer review.","tokens_in":11092,"tokens_out":2313,"would_cite":true,"duration_ms":25112,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that prompting an LLM reward designer in English instead of Hindi, Tamil, or Tulu changes both task performance and fairness in restless-bandit resource allocation.","keywords":["restless multi-armed bandits","reward function design","large language models","multilingual prompts","low-resource languages","fairness","demographic parity","resource allocation"],"falsifier":"Have native speakers back-translate the eight Hindi, Tamil, and Tulu prompts into English without seeing the originals, then rerun the DLM pipeline with the back-translated prompts; if acceptable-reward rates and fairness metrics match the English-prompt results, the claimed language effects collapse, while if the gaps persist under faithful translations, they are confirmed.","tokens_in":10145,"feed_emoji":"🌐","tokens_out":5262,"duration_ms":47121,"temperature":0.7,"pith_summary":"This paper tests what happens when the goal prompts that steer an LLM-designed reward function for restless multi-armed bandits are written in Hindi, Tamil, and Tulu instead of English. Using the DLM pipeline on a synthetic public-health resource-allocation environment, it finds that English prompts yield acceptable reward functions more often and produce allocations that match the user's stated priorities more closely. It also finds that semantically identical prompts with different phrasing lead to different allocations, and that more complex prompts degrade performance in every language, with English degrading less. On fairness, prompts in low-resource languages and prompts with more intended features are more likely to create demographic-parity violations on features the user did not intend to target. The paper's point is that language choice and prompt wording are not neutral: they change both the effectiveness and the fairness of automated resource allocation.","feed_headline":"English prompts beat Hindi, Tamil, Tulu on LLM-designed rewards","feed_subtitle":"Prompt language shifts both task success and fairness in restless-bandit health allocation.","key_machinery":"The central object is the DLM (Decision-Language Model) pipeline: an LLM is given a chain-of-thought prompt with a goal and is asked to write a Python reward function, then an evolutionary search with an LLM-based reflection step refines proposals, and a Whittle-index policy solves the resulting restless multi-armed bandit, a sequential resource-allocation model where each arm's state evolves regardless of whether it is pulled. The paper varies the language of the goal prompt (English, Hindi, Tamil, Tulu) and the complexity and phrasing of the prompt, and measures outcomes with two instruments: the rate of 'acceptable' reward functions that use all and only the intended features, and demographic-parity variance across feature buckets, where high variance on unintended features counts as unfairness. The work these parts do is to separate language effects from prompt-complexity and phrasing effects in a controlled synthetic environment.","core_discovery":"On the paper's own terms, the central discovery is that the DLM algorithm's LLM-proposed reward functions are significantly better when prompted in English than in Hindi, Tamil, or Tulu, and that this gap shows up both in task performance and in fairness. The intended reward features are identified correctly at a higher rate under English prompts, and the resulting Whittle-index allocations deviate less from what the prompt asks for. The paper further claims that the exact phrasing of a prompt alters allocations even when the meaning is unchanged, that explicit statements of the goal help, and that increasing prompt complexity hurts all languages but hurts lower-resource languages more. On the fairness side, low-resource languages and more complex prompts are both highly likely to create demographic-parity variance on unintended features, which the paper treats as unfairness. Because the system is aimed at grassroots public-health workers who may prefer local languages, the claim implies that deployment language alone could shift who gets calls.","pith_inferences":["A reasonable reading is that the performance gap may reflect the LLM's weaker command of Hindi, Tamil, and Tulu rather than any property of those languages; if so, improving multilingual instruction-following or adding a prompt-rewriting stage could close most of the gap.","The fairness result suggests a cheap testable safeguard: compute demographic-parity variance on all non-target features before deployment and reject reward functions that exceed a threshold on unintended dimensions.","One could extend the study by back-translating the translated prompts and comparing reward quality on the back-translations, which would separate translation fidelity from intrinsic language handling.","The collapse past roughly three intended features hints that the evolutionary search and reflection stage, not the LLM's language ability, may be the binding constraint for complex goals."],"forward_implications":["English prompts yield significantly higher rates of acceptable reward functions than Hindi, Tamil, or Tulu prompts across the tested goal prompts.","Semantically equivalent rephrasings produce noticeably different allocations, so prompt wording is a performance variable, not a nuisance.","As prompts involve more intended features, success rates drop for all languages, and English remains more stable; past roughly three intended features, performance collapses regardless of language.","Low-resource-language prompts and more complex prompts are more likely to produce demographic-parity variance on unintended features, meaning unfair allocations along dimensions the user did not specify.","For deployed public-health allocation, these gaps imply that outcomes would differ across language communities unless the DLM pipeline is modified."],"supporting_citations":[{"why":"Supplies the DLM algorithm under test, the LLM-to-reward-function pipeline with evolutionary search and reflection, and the public-health motivation.","marker":"Behari et al. 2024"},{"why":"Supplies the synthetic environment, feature set, transition-probability construction, and the choice of Gemini 1.0 Pro, and is the prior prioritization-strategy work this paper extends.","marker":"Verma et al. 2024"},{"why":"Supplies the Whittle-index policy used to solve the restless bandit once the LLM-designed reward functions are defined.","marker":"Whittle 1988"},{"why":"Supplies Gemini 1.0 Pro, the LLM that proposes and reflects on reward functions in the experiments.","marker":"Team et al. 2023"},{"why":"Supplies demographic parity, the fairness notion behind the DP-variance metric the paper uses to detect unfairness on unintended features.","marker":"Kusner et al. 2017"},{"why":"Supplies the restless-bandit transition model used to express the difference between active and passive transition probabilities as a function of feature values.","marker":"Liu, Liu, and Zhao 2013"}],"fun_headline_variants":["English prompts beat Hindi, Tamil, Tulu in LLM reward design","LLM reward fairness drops with low-resource prompts","English LLM prompts yield fairer restless-bandit rewards","Hindi, Tamil, Tulu prompts hurt LLM reward fairness","Prompt complexity worsens LLM reward gaps in low-resource languages"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Hindi, Tamil, and Tulu prompts are faithful translations of the English prompts; the paper provides no back-translation, native-speaker verification, or translation-quality scores, so the observed gaps could in principle be artifacts of translation quality rather than of language itself.","fun_headline_variants_meta":{"raw":{"variants":["English prompts beat Hindi, Tamil, Tulu in LLM reward design","LLM reward fairness drops with low-resource prompts","English LLM prompts yield fairer restless-bandit rewards","Hindi, Tamil, Tulu prompts hurt LLM reward fairness","Prompt complexity worsens LLM reward gaps in low-resource languages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001252,"raw_usage":{"total_tokens":5170,"prompt_tokens":1024,"completion_tokens":4146,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":4060}},"tokens_in":640,"tokens_out":4146,"duration_ms":31126,"temperature":1.0,"reasoning_tokens":4060,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:00:29.104527+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have native speakers back-translate the eight Hindi, Tamil, and Tulu prompts into English without seeing the originals, then rerun the DLM pipeline with the back-translated prompts; if acceptable-reward rates and fairness metrics match the English-prompt results, the claimed language effects collapse, while if the gaps persist under faithful translations, they are confirmed.","supporting_citations":[{"cited_title":"M.; and Tambe, M","cited_arxiv_id":null,"evidence_quote":"Supplies the DLM algorithm under test, the LLM-to-reward-function pipeline with evolutionary search and reflection, and the public-health motivation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the synthetic environment, feature set, transition-probability construction, and the choice of Gemini 1.0 Pro, and is the prior prioritization-strategy work this paper extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Whittle-index policy used to solve the restless bandit once the LLM-designed reward functions are defined."},{"cited_title":"J.; Loftus, J.; Russell, C.; and Silva, R","cited_arxiv_id":null,"evidence_quote":"Supplies demographic parity, the fairness notion behind the DP-variance metric the paper uses to detect unfairness on unintended features."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the restless-bandit transition model used to express the difference between active and passive transition probabilities as a function of feature values."}],"review_version":1}