{"id":"f8f3153e-d77f-4a48-adcc-2f512d6e26db","arxiv_id":"2501.12894","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A 30-user evaluation of an educational recommender with user control at input, process, and output found positive ratings and correlations among perceived control, transparency, trust, and satisfaction, but no baseline supports the causal claims.","lead":"This paper evaluates an educational recommender system in which learners can control the input, process, and output of recommendations. In a 30-student study, users rated the system positively and perceived control correlated with transparency, trust, and satisfaction, but the design had no comparison condition.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-condition design with no baseline cannot support the causal 'impact of user control' claims; the reported correlations are within-condition self-reports from the same questionnaire.","rationale":"The reader's verdict is sound. The paper is a useful design case and exploratory correlation study, but its headline claims use causal language ('positive impact,' 'leads to increased transparency') unsupported by the single-condition design. I agree with the reader's weakest assumption: the study assumes that a single post-task questionnaire from one group, with no baseline and no manipulation check, is enough to attribute the positive ratings and correlations to user control. This premise is load-bearing for the central claim, and it fails. The paper's own future-work section concedes the need for a more comprehensive study. The proposed between-subjects comparison would directly test whether user control, rather than the overall system or context, drives the observed effects. Therefore, the appropriate recommendation remains REJECT as submitted, with a clear path toward a controlled evaluation.","tokens_in":15376,"tokens_out":3048,"duration_ms":32517,"concrete_test":"Run a between-subjects experiment with the same CourseMapper ERS but two conditions: (a) full user control as in the paper, and (b) a matched system where the input, process, and output controls are present but non-functional or hidden, while recommendations are generated identically. Randomly assign roughly 30 new participants per condition, administer the same ResQue questionnaire plus a manipulation check (e.g., 'I could adjust how recommendations were generated'). If condition (b) yields statistically similar transparency, trust, and satisfaction means and correlations, the central causal claim is falsified; if condition (a) significantly outperforms (b), the claim gains support.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim that user control positively impacts transparency, trust, and satisfaction rests on a one-arm study. Section 4.1 describes a single group (N=30) using the ERS with all control features; there is no control condition without user control, no within-subjects manipulation, and no manipulation check verifying that participants actually perceived or used the controls differently. Consequently, the high means in Table 2 and the correlations in Figure 6 are consistent with many non-control explanations: the underlying recommender's quality, novelty of the MOOC platform, social desirability, demand characteristics from the guided tasks, or the demo video. RQ1 asks how complementing an ERS with user control impacts perceptions, but 'complementing' is never implemented as a comparison. RQ2's 'effects of user control' is tested by correlating a self-reported two-item control measure with other self-reported measures taken from the same questionnaire at the same time; this is a common-method correlation, not evidence that control features cause trust, transparency, or satisfaction. The paper itself acknowledges the limitation in Section 6 ('we plan to conduct a more comprehensive user study'), and Section 5.3's speculative explanations about cognitive load and trust underscore that the design cannot distinguish competing mechanisms. Thus the load-bearing premise—that observed positive ratings and correlations are attributable to user control—is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the design of an interactive educational recommender system (ERS) built into the MOOC platform CourseMapper, with user control at three levels: input (selecting and weighting misunderstood concepts), process (choosing a recommendation algorithm and ranking weights), and output (sorting, saving, and giving feedback on recommendations). The authors report a single-group online user study (N=30) in which participants performed guided tasks and then completed a post-task questionnaire based on the ResQue framework. The paper claims that user control has a positive impact on users' perceptions, that user control correlates strongly with transparency and moderately with trust and satisfaction, and that transparency and trust each correlate with satisfaction but less with each other.","tokens_in":15689,"tokens_out":2557,"duration_ms":26977,"significance":"If the causal claims were supported, the paper would make a useful contribution by systematically combining input-, process-, and output-level control in an educational recommender and jointly evaluating transparency, trust, and satisfaction. The UI design is informed by a literature review and iterative prototyping, and the evaluation uses a recognized framework (ResQue) with bootstrap confidence intervals and adjusted p-values. However, the evaluation design is a one-arm post-test study with no baseline or manipulation check, so the central claim that user control causes transparency, trust, or satisfaction is not supported by the data. The paper is nevertheless informative as a descriptive account of user perceptions of a control-rich ERS and as a design exploration, which may be of interest to practitioners.","major_comments":[{"comment":"The study uses a single condition: all 30 participants interacted with the full ERS containing all user-control features. There is no control condition without user control, no within-subjects manipulation, and no manipulation check. Consequently, RQ1 ('How does complementing an ERS with user control impact users' perceptions of the ERS?') cannot be answered from these data. The high mean scores in Table 2 and Figure 5 could be due to the recommender's underlying quality, the guided task structure, the demo video, or demand characteristics, rather than to the presence of user-control features. The causal language in the abstract and in Section 5.1 ('user control over the ERS can lead to relevant and novel recommendations') is therefore unsupported.","section":"Section 4.1 and RQ1 (Section 1)"},{"comment":"RQ2 asks about the 'effects of user control' on transparency, trust, and satisfaction, and Section 5.2 states that user control 'leads to increased transparency' and that 'user control over the RS positively influenced their satisfaction.' These claims rest on Pearson correlations between a self-reported two-item control measure and other self-reported measures collected in the same post-task questionnaire. This is a common-method correlation analysis, which cannot establish causal effects. The absence of any behavioral measure of control use or a manipulation check further weakens the inference. The paper itself acknowledges in Section 6 that 'we plan to conduct a more comprehensive user study,' which is consistent with the present design being insufficient for the causal claims made.","section":"Section 5.2 and Figure 6"},{"comment":"The paper explains the relatively low correlation between transparency and trust with speculative mechanisms such as cognitive load from too much control and the absence of explanations. However, no cognitive-load measures, open-ended responses, or explanation manipulations were collected, so these explanations are untestable in the current design. The text should be clearly framed as hypotheses for future work rather than as findings from this study.","section":"Section 5.3"}],"minor_comments":[{"comment":"The sentence 'Most of the participants were familiar with the use of RSs (n=24, 63%)' contains an inconsistency: 24 out of 30 is 80%, not 63%; the 63% figure applies to the following phrase about interacting with RSs (n=19, 63%).","section":"Section 4.1"},{"comment":"Several constructs (e.g., Interface Adequacy, Perceived Usefulness, Use Intentions) are measured with multiple items, but only a single mean and SD are reported for the whole construct. Reporting item-level statistics or reliability coefficients (e.g., Cronbach's alpha) would improve interpretability.","section":"Table 2"},{"comment":"The exact Pearson correlation coefficients and adjusted p-values are not stated in the text; the figure shows them visually, but numeric values in the text or a table would make the results more transparent and reproducible.","section":"Figure 6"},{"comment":"There is a typo in the list of participant countries: '2 Chineese' should be '2 Chinese.'","section":"Section 4.1"}],"recommendation":"reject","confidential_remarks":"This appears to be a workshop-style paper (IntRS'24). As a short demo or system description it could be acceptable, but as a full research article the absence of any baseline or control condition is a load-bearing flaw that cannot be repaired by local revisions; the research questions would need to be reformulated and new data collected. The paper's own future-work statement acknowledges this gap."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The design work is the real contribution here; the causal claims about user control are not supported by the evaluation. The paper gives a genuinely systematic account of how to add input, process, and output control to an educational recommender, grounded in a broad review of interactive recommender UI mechanisms, and it studies transparency, trust, and satisfaction together, which has not been done in the ERS literature. That is a legitimate contribution and worth keeping.\n\nWhere the paper breaks down is in the inference from data to conclusions. The user study is a single condition with N=30, no baseline, no within-subject manipulation, and no manipulation check. Participants all used the same system with all controls available, then filled one post-task questionnaire. So the high means in Table 2 and the correlations in Figure 6 cannot be attributed to user control specifically. The positive ratings could come from the underlying recommender, the novelty of the platform, the guided tasks, or social desirability. The paper's own Section 6 concedes that a more comprehensive study is needed, and Section 5.3 is openly speculative about the trust/transparency gap. The abstract's 'positive impact' and RQ1's 'impact' language overshoots what a one-arm correlation design can show. The title also promises different levels of user control, but the analysis treats control as a single self-report score.\n\nThe correlations for RQ2 and RQ3 are underreported: no raw data, no exact r values in the text, and the measures are single- or two-item self-reports from the same questionnaire at the same time. That makes common-method variance a real alternative explanation for the 'transparency through controllability' link. These are not minor nits; they bear on the central claim.\n\nI would not want to see this paper rejected for the design work or the related work, which are solid. But as submitted, the causal framing should not stand. The path forward is either a controlled or within-subjects comparison with a baseline condition and a manipulation check, or an honest reframing of the paper as an exploratory design case with correlational findings. Given the quality of the design section and the novelty of the combined outcome measures, I would still send it to a serious peer review. For my own work, I would not cite it for the empirical claims, but the UI design analysis is worth a look. This is a useful read for anyone building interactive ERS interfaces, and a useful caution for anyone evaluating them.","headline":"A solid design contribution with an evaluation that cannot support the causal claims; the paper should be revised toward an exploratory design case.","tokens_in":16122,"tokens_out":3350,"would_cite":false,"duration_ms":34221,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"User control in an educational recommender is associated with higher transparency, trust, and satisfaction, a 30-user study finds.","keywords":["educational recommender systems","interactive recommender systems","user control","transparency","trust","satisfaction","MOOC platform","user study"],"falsifier":"A controlled experiment that runs the same educational recommender with the control widgets active for one group and hidden for another, measuring the same questionnaire items, would falsify the causal claim if the control group shows no significantly higher transparency, trust, or satisfaction scores.","tokens_in":15135,"feed_emoji":"🎓","tokens_out":7212,"duration_ms":61701,"temperature":0.7,"pith_summary":"This paper claims that letting learners control an educational recommender system at the level of the input profile, the recommendation algorithm, and the output list improves how users perceive the system. The authors designed such controls into the ERS module of the CourseMapper MOOC platform and evaluated them in an online study with 30 students. They report positive ratings across accuracy, novelty, interaction adequacy, ease of use, transparency, trust, satisfaction, and use intentions. They further report that user control correlates strongly with transparency and moderately with trust and satisfaction. The authors conclude that user control is a viable design lever for making educational recommenders more transparent and satisfying, while noting that transparency and trust should be evaluated as distinct goals.","feed_headline":"User control in learning recommenders tracks transparency and trust","feed_subtitle":"A 30-user study finds three-level control widgets correlate with higher transparency, trust, and satisfaction.","key_machinery":"The carrying mechanism is a set of interactive control widgets organized along the three standard levels of a recommender. At the input level, learners select \"did not understand\" concepts, adjust concept weights with sliders, and include or exclude concepts with checkboxes. At the process level, they choose among four recommendation algorithms via radio buttons and adjust ranking-factor weights with sliders whose effects appear in real-time progress bars. At the output level, they mark recommendations as helpful or not helpful, with a follow-up selection of which concepts were clarified, sort by similarity, recency, or views, and save items for later. This three-level control design is evaluated with a post-task questionnaire based on a standard user-centric evaluation framework, and the paper's statistical claims rest on Pearson correlations with bootstrap confidence intervals.","core_discovery":"The central discovery is that user control at all three levels of a recommender, input, process, and output, is associated with positive user-perceived benefits in an educational setting, and specifically that user control strongly correlates with transparency, moderately with trust, and moderately with satisfaction. In the same data, transparency moderately correlates with satisfaction, trust strongly correlates with satisfaction, but transparency and trust are less correlated with each other. The authors interpret this as evidence for \"transparency through controllability\": users understand why items are recommended because they can shape the profile, algorithm, and ranking themselves. They also argue that because transparency and trust move somewhat independently, evaluations of interactive educational recommenders should treat them as separate constructs.","pith_inferences":["A natural next test, not run in this paper, would be to compare the three control levels separately to see whether input, process, or output control drives most of the transparency gain.","If transparency through controllability is real, adding explanation text alongside the control widgets should push trust up further; the authors themselves hypothesize this as transparency through explanation.","A baseline condition without any control widgets would be needed to separate the effect of control from a general novelty or interface-quality effect; the current design has no such comparison."],"forward_implications":["Designers of educational recommender systems can treat controllability as a practical route to transparency, since control and transparency showed the strongest correlation in the study.","Because trust strongly correlates with satisfaction, improving either one is likely to carry the other upward.","Transparency and trust should be measured as separate evaluation dimensions in interactive educational recommenders, since they correlated only weakly with each other.","Providing control at all three levels, input, process, and output, is feasible inside a MOOC platform and was rated positively by learners with varied backgrounds.","The authors' account implies a trade-off: too much control can raise cognitive load, which may explain why transparency and trust do not always move together."],"supporting_citations":[{"why":"It defines the three levels of user control (input, process, output) and frames the interactive recommender design space.","marker":"[8]"},{"why":"It supplies the user-centric evaluation questionnaire that all reported measures are taken from.","marker":"[46]"},{"why":"It describes the personal-knowledge-graph-based recommendation algorithms that the process-level controls select and weight.","marker":"[33]"},{"why":"It introduces the MOOC platform in which the educational recommender module is embedded.","marker":"[32]"},{"why":"It defines the recommendation goals of transparency, trust, and satisfaction and links explanations to these goals.","marker":"[31]"},{"why":"It documents the cognitive-load trade-off of additional controls, which the authors use to explain the weaker transparency-trust link.","marker":"[7]"},{"why":"It establishes that controllability and explainability affect transparency in social recommenders, a direct predecessor of the transparency-through-controllability idea.","marker":"[25]"},{"why":"It shows that control and transparency can be provided together in a social recommender, another load-bearing prior result.","marker":"[19]"}],"fun_headline_variants":["User control boosts trust, transparency in learning recommenders","More user control, more trust in learning recommenders","Three-level user control lifts recommender transparency and trust","Control levels in recommenders predict transparency, trust"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that a single post-task questionnaire from 30 users, with no baseline system and no check that participants actually understood or used the control features, is enough to attribute the positive ratings and correlations to the user control design.","fun_headline_variants_meta":{"raw":{"variants":["User control boosts trust, transparency in learning recommenders","More user control, more trust in learning recommenders","Three-level user control lifts recommender transparency and trust","Control levels in recommenders predict transparency, trust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001285,"raw_usage":{"total_tokens":5239,"prompt_tokens":920,"completion_tokens":4319,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":4257}},"tokens_in":536,"tokens_out":4319,"duration_ms":33391,"temperature":1.0,"reasoning_tokens":4257,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T16:39:49.906246+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled experiment that runs the same educational recommender with the control widgets active for one group and hidden for another, measuring the same questionnaire items, would falsify the causal claim if the control group shows no significantly higher transparency, trust, or satisfaction scores.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It defines the three levels of user control (input, process, output) and frames the interactive recommender design space."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the user-centric evaluation questionnaire that all reported measures are taken from."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It describes the personal-knowledge-graph-based recommendation algorithms that the process-level controls select and weight."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It introduces the MOOC platform in which the educational recommender module is embedded."},{"cited_title":"Tintarev, J","cited_arxiv_id":null,"evidence_quote":"It defines the recommendation goals of transparency, trust, and satisfaction and links explanations to these goals."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It documents the cognitive-load trade-off of additional controls, which the authors use to explain the weaker transparency-trust link."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It establishes that controllability and explainability affect transparency in social recommenders, a direct predecessor of the transparency-through-controllability idea."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It shows that control and transparency can be provided together in a social recommender, another load-bearing prior result."}],"review_version":1}