{"id":"c5cb4612-62fc-42e3-9e29-20583d9c438c","arxiv_id":"2607.00211","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"Analysis of student-GenAI co-programming dialogues shows 78.8% lack mastery-oriented epistemic aims and reliable processes, while only 11.1% demonstrate high epistemic engagement.","lead":"The paper introduces Epistemic AI Literacy as a framework for analyzing student goals and thinking processes during generative AI use in programming tasks. Smart generalists might read it to understand how to measure and improve whether AI tools promote deep learning or just quick answers.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Validity of coding observable dialogue dimensions to AIR constructs remains unverified in text data","rationale":"The reader’s weakest_assumption correctly isolates the measurement-validity step that the empirical claim depends on. Because the provided abstract supplies only the final percentages and the full text (per the prompt) is the source for checking operationalization details, the same concern remains the single most load-bearing one; no other internal inconsistency is visible from the given material.","tokens_in":1809,"tokens_out":337,"duration_ms":16914,"concrete_test":"Report inter-rater reliability (Cohen’s kappa or Krippendorff’s alpha) on a held-out sample of at least 200 coded turns for all six dimensions; if any dimension falls below 0.65, recompute the 78.8% / 11.1% figures on the subset of reliably coded interactions and check whether the headline percentages shift by more than 10 points.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The headline percentages (78.8% low EAIL, 11.1% high engagement) are produced by classifying interactions on mastery-oriented aims plus the five process dimensions. For these counts to support the claim of prevalent lack of EAIL, the mapping from surface features (e.g., “outsourcing”, “epistemic justification”) to the AIR definitions must be both valid and low-error. Text-only transcripts can lose intent, shared context, or non-verbal cues; the paper identifies the dimensions but the load-bearing step is whether the subsequent annotation procedure demonstrably recovers the intended constructs without substantial misclassification.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces the Epistemic AI Literacy (EAIL) framework by adapting the AIR (aims, ideals, reliable processes) model to student-GenAI co-programming interactions. It identifies observable dimensions (mastery-oriented aims; outsourcing, explanation seeking, verification seeking, prompt monitoring, epistemic justification) in a large dialogue dataset and reports that 78.8% of interactions exhibit low EAIL (non-mastery aims plus less reliable strategies) while only 11.1% exhibit high epistemic engagement (mastery aims coupled with advanced strategies such as epistemic justification).","tokens_in":1941,"tokens_out":444,"duration_ms":13501,"significance":"If the annotation procedure is shown to be valid and reliable, the work supplies a process-oriented lens on AI literacy that could guide both assessment and instructional design in programming education. The use of an established external framework (AIR) on interaction data is a strength, but the absence of dataset size, collection protocol, coding details, and reliability metrics prevents evaluation of whether the headline percentages are reproducible or generalizable.","major_comments":[{"comment":"Abstract and results section: the claims of 78.8% low EAIL and 11.1% high engagement are presented without any information on total interactions analyzed, sampling method, annotation protocol, inter-rater reliability, or statistical tests. These omissions are load-bearing because the percentages constitute the central empirical claim.","section":"Abstract / Results"},{"comment":"Methods / coding procedure: the mapping from surface dialogue features (e.g., \"outsourcing\", \"epistemic justification\") to AIR constructs is asserted but not validated against text-only data; no evidence is supplied that the annotation recovers intended epistemic aims and processes without substantial context loss or misclassification.","section":"Methods"}],"minor_comments":[{"comment":"The term \"large dialogue dataset\" is used without a citation or size; a reference to the source corpus and its scale would improve reproducibility.","section":"Abstract"},{"comment":"Notation for the five process dimensions is introduced without a table or explicit operational definitions; a summary table would aid clarity.","section":"Framework section"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback emphasizing the need for methodological transparency. We agree that the submitted manuscript lacks essential details on the dataset and annotation procedure, which are required to substantiate the central empirical claims, and we will revise the paper to address these points.","responses":[{"response":"We agree that these omissions are significant. The current manuscript does not report the total number of interactions analyzed, sampling method, annotation protocol, inter-rater reliability, or statistical tests. In the revised version, we will add this information to the Methods section (including dataset size, collection protocol, coding details, and reliability metrics such as inter-rater agreement) and update the abstract and results sections to reference these details, enabling evaluation of the reported percentages.","revision_made":"yes","referee_comment":"[Abstract / Results] Abstract and results section: the claims of 78.8% low EAIL and 11.1% high engagement are presented without any information on total interactions analyzed, sampling method, annotation protocol, inter-rater reliability, or statistical tests. These omissions are load-bearing because the percentages constitute the central empirical claim."},{"response":"The manuscript asserts the mapping of dialogue features to AIR constructs without providing validation evidence or reliability metrics for text-only annotation. We will revise the Methods section to include a detailed description of the coding scheme development, explicit mappings with examples, and inter-rater reliability statistics. A comprehensive external validation study to rule out context loss may exceed the scope of the current work, but we will add available evidence from the annotation process and acknowledge limitations.","revision_made":"partial","referee_comment":"[Methods] Methods / coding procedure: the mapping from surface dialogue features (e.g., \"outsourcing\", \"epistemic justification\") to AIR constructs is asserted but not validated against text-only data; no evidence is supplied that the annotation recovers intended epistemic aims and processes without substantial context loss or misclassification."}],"tokens_in":1409,"tokens_out":426,"duration_ms":26423,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this work builds Epistemic AI Literacy on top of the AIR framework and applies it to co-programming dialogues, finding that 78.8% of interactions show low engagement while only 11.1% reach high levels with mastery aims plus epistemic justification.\n\nWhat is actually new is the list of detectable markers: mastery-oriented aims plus the five process dimensions (outsourcing, explanation seeking, verification seeking, prompt monitoring, epistemic justification). These give a concrete way to tag interaction transcripts rather than relying on abstract definitions alone.\n\nThe paper does a straightforward job of extending an existing model to GenAI use without claiming to reinvent the base theory. The dimensions are chosen to be visible in text, which is a practical move for anyone who wants to analyze large dialogue sets.\n\nThe soft spot is the empirical claim. The percentages come from classifying interactions on those dimensions, yet the abstract supplies no dataset size, collection details, coding protocol, or reliability numbers. Text transcripts routinely drop intent and shared context, so the mapping from surface features to AIR constructs can introduce error. The stress-test note is right to flag this as the step that needs checking; without validation examples or agreement stats, the 78.8% and 11.1% figures stay hard to trust.\n\nThis is for researchers in educational technology and AI literacy who care about the quality of student thinking rather than raw usage counts. Someone building analysis tools or classroom interventions could borrow the dimension list as a starting point.\n\nIt deserves peer review so referees can see the full methods and data handling. The framework idea is workable if the coding side holds up.","headline":"The paper defines observable dimensions for epistemic engagement in student-GenAI programming but the headline percentages rest on unverified coding from text data.","tokens_in":2404,"tokens_out":406,"would_cite":false,"duration_ms":25960,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Most student-GenAI programming interactions rely on non-mastery aims and less reliable strategies like outsourcing.","keywords":["epistemic AI literacy","generative AI","student-AI interaction","epistemic aims","AIR framework","programming education","dialogue analysis","co-programming"],"falsifier":"A fresh round of independent annotation on the same dialogue dataset that produces markedly different shares of mastery-oriented aims or high-epistemic-engagement turns.","tokens_in":2689,"feed_emoji":"🧠","tokens_out":671,"duration_ms":31959,"temperature":0.7,"pith_summary":"The paper introduces Epistemic AI Literacy as a process-oriented view of how students build and evaluate knowledge while using generative AI for programming. It applies the AIR framework to label dialogue turns for aims such as mastery orientation and for processes such as outsourcing, verification seeking, and epistemic justification. Analysis of a large set of student-AI exchanges shows that nearly 79 percent of interactions use non-mastery aims paired with less reliable moves, while only about 11 percent reach high engagement through mastery aims combined with justification. A reader would care because the pattern suggests current everyday use of AI tools may not automatically produce thoughtful problem-solving habits in learners.","feed_headline":"78% of student-AI programming chats skip mastery aims","feed_subtitle":"Dialogue study finds most interactions use outsourcing and verification rather than justification or deep knowledge goals.","key_machinery":"Epistemic AI Literacy (EAIL), which detects mastery-oriented aims and reliable epistemic processes such as epistemic justification within human-AI dialogue data by applying the AIR framework.","core_discovery":"Drawing on the AIR framework, the study identifies observable dimensions of epistemic aims (mastery-oriented aims) and epistemic processes (outsourcing, explanation seeking, verification seeking, prompt monitoring, and epistemic justification) in GenAI-supported co-programming dialogues. The results show that 78.8 percent of student-GenAI interactions rely on non-mastery-oriented aims and less reliable epistemic strategies, whereas only 11.1 percent of interactions display high epistemic engagement in which mastery-oriented aims are coupled with epistemic justification in a more reliable epistemic process.","pith_inferences":["Educational designs that prompt students to state justifications may raise the share of high-engagement turns.","GenAI interfaces for classrooms could add cues that favor explanation seeking over direct outsourcing.","Differences in EAIL levels across interactions may correspond to different long-term gains in programming competence."],"forward_implications":["Most student-GenAI co-programming interactions operate with non-mastery aims and less reliable strategies such as outsourcing and verification seeking.","A small share of interactions reach high epistemic engagement when mastery aims pair with epistemic justification.","Observable dimensions of aims and processes can be identified and scaled across large dialogue datasets from human-AI co-programming.","Epistemic AI literacy can be treated as an emergent property of dynamic interactions rather than a static skill set."],"fun_headline_variants":["78% student-AI chats skip mastery epistemic aims","Only 11% high engagement in GenAI co-programming","Outsourcing rules 79% of student-AI coding talks","Weak epistemic processes in most AI student sessions"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The observable dimensions extracted from text dialogues accurately capture the underlying epistemic aims and processes defined by the AIR framework.","fun_headline_variants_meta":{"raw":{"variants":["78% student-AI chats skip mastery epistemic aims","Only 11% high engagement in GenAI co-programming","Outsourcing rules 79% of student-AI coding talks","Weak epistemic processes in most AI student sessions"]},"model":"grok-4.3","cost_usd":0.004657,"raw_usage":{"total_tokens":2328,"prompt_tokens":715,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":46574500,"prompt_tokens_details":{"text_tokens":715,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1551,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":715,"tokens_out":62,"duration_ms":14155,"temperature":1.0,"reasoning_tokens":1551,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-02T19:01:02.462504+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A fresh round of independent annotation on the same dialogue dataset that produces markedly different shares of mastery-oriented aims or high-epistemic-engagement turns.","supporting_citations":[],"review_version":1}