{"id":"14721a46-7175-4915-be25-c684451e487f","arxiv_id":"2606.29251","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"LLM compression of filings and earnings calls often changes the source-implied bear/neutral/bull decision; agentic multi-candidate auditing against the source reduces those flips.","lead":"LLM summaries of financial filings can flip investment judgments even when they stay fluent and factually plausible. The paper defines decision-preserving information fidelity, diagnoses decontextualization and model dependency, and shows multi-candidate source-audited compression reduces those flips.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The fidelity metric is defined by the same LLM family that may share compressor biases, so flips may partly measure judge-compressor alignment rather than source-decision preservation.","rationale":"The reader correctly isolates the single weakest link: fidelity is defined as preservation of the decision induced by the source, yet that decision is measured only by one LLM proxy without human or market grounding. Self-agreement (Table 1) and the reread noise floor rule out pure stochasticity of E, and the decontextualization add-back (Figure 4) plus multi-compressor disagreement (Figures 2–3, 6) supply internal diagnostics that something systematic is happening. Those diagnostics do not, however, establish that the systematic thing is loss of the investment judgment a human would form from the source. My concern is therefore the same as the reader’s weakest_assumption, sharpened only by noting possible shared LLM bias between compressors and E. Because the paper already states this limitation and the experimental design is otherwise clear and multi-baseline, the appropriate stance remains CONDITIONAL rather than REJECT: the contribution is accept-shaped once the proxy is treated as provisional and independent judges (or human labels) are supplied. No stronger internal inconsistency appears in the reported equations or tables.","tokens_in":14040,"tokens_out":643,"duration_ms":6890,"concrete_test":"Re-run the full Table 2 pipeline on the same 597 sources with at least two held-out decision models of different families (e.g., GPT-class and a finance-tuned open model) under identical D.5 prompts, plus a small human-analyst subset (N≥50) labeling source vs. compressed bear/neutral/bull. If Flip/TVD rankings reverse or ACC’s reduction vs. naive disappears under the alternative judges, the proxy assumption fails and the fidelity claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim treats Flip and TVD (Eqs. 4–5) as source-relative decision preservation: compression loses fidelity when it changes the decision induced by the source. That definition is operationalized solely by one fixed decision model E = Gemini-3.1-Flash-Lite producing a three-class next-quarter regime belief (Section 3.1, Decision Model Prompt D.5). Reliability checks (Table 1 self-agreement / Fleiss’ κ) only show that E is stable when rereading the same text; they do not show that E’s top label matches human investment judgment or that E is independent of the compressors. Because compressors include other frontier LLMs and the decision prompt itself is an LLM judgment task, a flip can arise when the compressor and E disagree on which caveats matter, even if a human (or a different judge) would still call the summary faithful. The industry IC lift is directionally consistent but non-reproducible and still uses the same product stack. Limitations already flag absence of human analysts and market outcomes; that gap is load-bearing for the claim that the measured flips are information-fidelity losses rather than judge-model idiosyncrasy.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that LLM context compression of financial disclosures can change the investment judgment supported by the original source even when the compressed text remains fluent and factually plausible. It defines information fidelity as source-relative decision preservation, operationalized by Decision Flip rate and total variation distance (TVD) between a fixed decision model’s bear/neutral/bull beliefs on the full source versus the compressed text (Eqs. 2–5). On S&P 100 10-Q MD&A and earnings-call transcripts, one-shot compressors exceed a rereading noise floor; the authors diagnose decontextualization (selective loss of context facts) and model dependency (compressor-specific decision tilts), and propose Agentic Context Compression (ACC), which generates multiple contextualized candidates and audits disagreements against source spans. Under a fixed 20-bullet budget, ACC reduces flips relative to naive and multi-LLM baselines (Table 2), with supporting diagnostics in Figures 2–4 and a commercial forecasting case study.","tokens_in":14377,"tokens_out":1419,"duration_ms":19681,"significance":"If the measured flips largely reflect genuine loss of decision-relevant context rather than judge idiosyncrasy, the paper supplies a useful evaluation criterion for financial and agentic compression that goes beyond reconstruction, factuality, or token efficiency. The experimental design is careful on several dimensions: fixed bullet budget, explicit rereading noise floor, decision-model reliability (Table 1), multiple compressors and baselines including token pruning and an integrator, a fact-role inventory with add-back recovery (Figure 4), and budget-sensitivity checks (Appendix A). ACC is a concrete, source-grounded procedure rather than a purely prompt-level tweak. These elements make the work a credible contribution to high-stakes LLM evaluation and agentic context engineering, provided the decision-proxy assumption is strengthened.","major_comments":[{"comment":"Section 3.1 and Eqs. (2)–(5) define information fidelity solely via one decision model E (Gemini-3.1-Flash-Lite) over a three-class next-quarter return regime (prompt D.5). Table 1 establishes re-read stability (self-agreement, Fleiss’ κ), not that E’s top label matches human investment judgment or is independent of compressor families. Because several compressors are frontier LLMs and one is Gemini-family, flips and TVD can partly measure judge–compressor disagreement about which caveats matter rather than source-decision preservation. This is load-bearing for the central claim and for interpreting Table 2 and Figures 2–3. At minimum the paper should (i) re-evaluate a substantial subset with at least one independent judge model and, if feasible, a small human-analyst panel, and (ii) report flip/TVD agreement across judges; otherwise claims should be scoped explicitly as “decision-model","section":null},{"comment":"Section 5 reports that ACC improved forecasting IC by 8.3% over the original-source baseline and 23.8% over naive compression while cutting source-relative flips by 59.5%, with absolute product metrics omitted for confidentiality. These numbers are used to argue real-world relevance of fidelity, but they are non-reproducible and still sit inside a commercial stack that may share modeling choices with the paper’s decision framing. Either provide a reproducible public proxy (e.g., open IC-style evaluation on the same S&P 100 panel with a fixed forecasting head) or move the case study to a clearly labeled anecdotal appendix and avoid quantitative claims that cannot be audited.","section":null},{"comment":"Figure 4B’s add-back diagnostic is important evidence that missing context facts drive flips, but restored context exceeds the fixed budget B that defines the compression task (Eq. 1). The main ACC gains in Table 2 are therefore only partially explained by the decontextualization mechanism under the same constraint that the method must satisfy. The paper should either (a) run a budget-respecting recovery/selection ablation (e.g., swap in context facts while dropping lower-value bullets) or (b) clearly separate the diagnostic from the deployable claim and quantify how much of ACC’s flip reduction is attributable to better context retention within B versus multi-model adjudication.","section":null}],"minor_comments":[{"comment":"Figure 2 caption and text refer to a “noise floor” of 11.0%/8.8% flips, while Table 2’s “No compression” row reports the same quantities with TVD; keep notation and naming identical across figure and table.","section":null},{"comment":"In Section 3.2, “Contextualization” is described as revisiting the source when candidates miss interpretive detail, but the boxed prompt D.2 is a single-pass instruction; clarify whether multi-pass selection is actually implemented or only aspirational.","section":null},{"comment":"Appendix B reports mean off-diagonal inter-compressor agreement of 0.75; stating N and whether agreement is on induced decisions after compression (as the caption suggests) versus on raw labels would help readers interpret model dependency.","section":null},{"comment":"Cost columns in Table 2 are useful but units and pricing assumptions are unspecified; a short note on model APIs and date would improve reproducibility of the efficiency comparison.","section":null},{"comment":"Typos/style: abstract and introduction use “information fidelity” consistently, but occasional phrasing such as “its loss” (Section 3.1) has an unclear antecedent; a light copy-edit pass would help.","section":null}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid empirical methods contribution for a top AI/finance venue if the decision-proxy issue is addressed; without multi-judge or human validation I would be uncomfortable treating Flip/TVD as measuring “investment judgment” rather than a single LLM’s regime belief. The industry IC numbers are the weakest part for a serious journal and risk looking like marketing unless made auditable. Novelty is real relative to pure factuality/compression metrics, but the citation set already covers related summarization-distortion and financial LLM bias work reasonably well."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that this is a clean, usable measurement paper, not a slogan. On S&P 100 10-Q MD&A and earnings calls they show one-shot LLM compression flips a fixed decision model’s bear/neutral/bull call well above the reread noise floor, with TVD moving in parallel. They then give two diagnostics that actually explain the flips—selective stripping of context-role facts (caveats, offsets, expectation frames) under a fixed bullet budget, and compressor-specific directional tilts—and a practical fix, ACC, that generates contextualized candidates and greps the source on disagreements. Table 2 is the core: under the same ~4% token budget, ACC beats naive, chunking, LLMLingua variants, and a simple multi-model integrator on flip and TVD. The industry IC note is directional only, but it is consistent with the lab result.\n\nWhat is new is not “summaries can bias,” which the related-work section already owns. It is the decision-preservation criterion (flip + source-relative TVD), the fact-role inventory with add-back recovery, and the source-auditing multi-candidate procedure. Math is light but honest: fixed B, averaged decision runs, explicit noise floor, Fleiss’ κ on the judge. Citations are in the right places (decontextualization, prompt compression, financial LLM bias). No circular construction—ACC scores candidates on source spans, not on the decision labels.\n\nThe soft spot the stress-test flags is real and already half-admitted in Limitations: fidelity is defined by one Gemini flash-lite model on a three-class next-quarter regime, not by humans or returns. Self-agreement only shows the judge is stable, not that it is the right proxy or independent of compressor family. So some flips may be judge–compressor disagreement about which caveats matter. That does not erase the measurement; it bounds the claim. Budget sensitivity and inter-model agreement appendices are useful and do not hide residual flips.\n\nThis is for people building financial RAG/agent stacks and anyone evaluating long-context compression in high-stakes settings. I would bring it to reading group, cite the fidelity framing and the decontextualization diagnostic, and send it to peer review. Ask referees for a second judge family or a small human study; do not desk-reject.","headline":"Solid empirical paper: compression can flip LLM investment judgments above a reread floor, and multi-candidate source auditing helps; the load-bearing soft spot is the single LLM judge, not the experimental craft.","tokens_in":15044,"tokens_out":579,"would_cite":true,"duration_ms":5621,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"LLM compression of financial filings can flip investment decisions even when the short text stays fluent and factually plausible.","keywords":["information fidelity","context compression","LLM summarization","financial analysis","decontextualization","agentic systems","decision preservation","earnings calls"],"falsifier":"Have independent human analysts (or a held-out panel of decision models) read the same original sources and the same compressed bullets under the fixed budget: if excess decision flips vanish relative to rereading noise, or if source-audited multi-candidate selection no longer reduces those flips, the measured fidelity claim fails.","tokens_in":14929,"feed_emoji":"📊","tokens_out":667,"duration_ms":18089,"temperature":0.7,"pith_summary":"Financial decision-makers cannot read every filing and earnings call in full, so they rely on compressed context. This paper shows that when large language models produce those short contexts, a downstream decision model can reach a different bear, neutral, or bull judgment than the original source supports. The authors frame the failure as information-fidelity loss: the summary need not hallucinate to change the decision; it can keep headline facts while dropping the caveats, offsets, and framing that give those facts their investment meaning. They document two recurring patterns—decontextualization and compressor-specific tilts—and propose Agentic Context Compression, which generates multiple contextualized candidates and audits their disagreements against the source. The stake is practical: in finance and multi-step agent pipelines, efficient, fluent compression can still quietly rewrite the decision the source would have induced.","feed_headline":"LLM finance summaries can flip the investment call","feed_subtitle":"Fluent, fact-plausible compressions still change bear/bull judgments; source-audited candidates cut the flips.","key_machinery":"Information fidelity: a compression has high fidelity when the decision-model belief it induces stays close to the belief induced by the original source, measured by top-decision flip rate and total variation distance under a fixed budget. Agentic Context Compression carries the fix by generating multiple contextualized candidates, grepping the source for disagreements, and selecting the intact candidate with lower overclaim risk and better source-locus coverage.","core_discovery":"On real S&P 100 quarterly MD&A sections and earnings-call transcripts, one-shot LLM compression under a fixed bullet budget more than doubles decision-flip rates above a no-compression rereading floor, while also shifting belief distributions. The loss is not only random degradation: compressors drop context facts disproportionately and different models tilt the same source in different directions. Agentic Context Compression, which generates multiple contextualized candidates and audits disagreements against short source spans, reduces those flips more than naive prompting, chunking, token pruning, or simple multi-model merging under the same budget.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLM finance compressions double decision flips vs rereading","Fluent LLM summaries still reverse bull/bear investment calls","Compressors drop caveats and tilt filings model by model","Agentic multi-candidate audit cuts fidelity loss under budget","Decontextualized evidence alone can flip the investment judgment"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The paper treats flips and belief shifts from one three-class LLM decision model as a faithful stand-in for real investment judgment, without human analysts or market outcomes.","fun_headline_variants_meta":{"raw":{"variants":["LLM finance compressions double decision flips vs rereading","Fluent LLM summaries still reverse bull/bear investment calls","Compressors drop caveats and tilt filings model by model","Agentic multi-candidate audit cuts fidelity loss under budget","Decontextualized evidence alone can flip the investment judgment"]},"model":"grok-4.5","effort":"low","cost_usd":0.004548,"raw_usage":{"total_tokens":1282,"prompt_tokens":786,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":45480000,"prompt_tokens_details":{"text_tokens":786,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":434,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":786,"tokens_out":62,"duration_ms":3588,"temperature":1.0,"reasoning_tokens":434,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T11:01:21.022297+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Have independent human analysts (or a held-out panel of decision models) read the same original sources and the same compressed bullets under the fixed budget: if excess decision flips vanish relative to rereading noise, or if source-audited multi-candidate selection no longer reduces those flips, the measured fidelity claim fails.","supporting_citations":[],"review_version":2}