{"id":"34919f97-38ef-422e-8dc5-54081fd4b7cb","arxiv_id":"2608.00617","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Text-to-video generators under-develop irreversible processes: near-zero progress and 92-100% stasis versus real footage (rho=+0.40, 35% static), confirmed by human raters.","lead":"This paper measures whether AI-generated videos show one-way changes such as rusting or melting, and finds a reliable way to tell. The main finding is that video generators barely develop irreversible processes, producing near-static clips instead of advancing them, unlike real time-lapse footage.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CLIP directional readout may measure general motion rather than irreversible-attribute progress; the seven-model separation lacks a non-attribute control.","rationale":"The reader's weakest_assumption names the same risk: the per-category CLIP directional readout may be confounded by overall motion, lighting, or quality, and the human 0–4 rating may track dynamicness. My reading of the paper supports this as the single most load-bearing concern. The paper's own generic metric shows generated clips are globally less dynamic, so without a non-attribute control the reported attribute-progress gap cannot be attributed specifically to irreversible-process development. However, the paper is careful to claim under-development rather than systematic reversal, and it provides human validation and category-level significance; the concern weakens the specificity of the diagnosis but does not overturn the core observable (generated clips barely change on the attribute axis while real clips do). The reader's CONDITIONAL verdict already reflects this fragility, and my proposed control would sharpen or resolve it. I therefore recommend no change to the verdict.","tokens_in":23672,"tokens_out":4238,"duration_ms":56775,"concrete_test":"Run the identical progress/stasis protocol on the same 108 real + ~560 generated clips with a control CLIP direction that is attribute-irrelevant, e.g. a random unit vector u_rand or a reversible axis such as 'brightness increasing'. If real footage again shows ρ≈+0.40 and low stasis on the control axis, the reported separation is a general-dynamicness artifact. A complementary check: temporally shuffle the frames of real clips; if shuffled real footage retains high progress, the readout is not capturing temporal attribute development.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central separation (Sec. III-l, Table III) — real footage ρ=+0.40, 35% stasis vs. generators ρ≈0, 92–100% stasis — is obtained with a per-category CLIP directional readout a(x)=⟨φ(x),u⟩. The paper itself shows (Sec. III-f) that generated clips move 2–3× less overall, so this readout may be tracking general dynamicness rather than the specific irreversible attribute. The only direct validity check (Sec. III-m) is a 0–4 human rating of 'how much the process advances' on 60 clips; that rating can itself reflect general motion, and the paper reports no control separating attribute progress from any directional change (lighting, camera, quality). If the readout responds to non-attribute motion, the real/generated gap in ρ and stasis would shrink or vanish once motion is controlled, and the 'under-development of irreversible processes' claim would reduce to 'generated clips are static.' This is the load-bearing assumption: a(x) is a faithful, attribute-specific progress measure, invariant to non-attribute visual change. It is asserted rather than established for the large-scale diagnosis, and the independent chroma probe used in the guidance experiments is not applied to the seven-model study.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses whether text-to-video generators respect irreversible physical processes. It argues that naive metrics—a per-clip violation rate, a variance-normalized reversal residual, and a generic monotonicity score—are null-degenerate or reward stasis, and it proposes a two-part protocol: progress (correlation between time and an attribute-specific readout) and a stasis rate read against a matched real baseline. Applying this protocol to seven T2V models and 108 real reference clips, the paper reports a clean separation: real footage advances (rho=+0.40, 35% static) whereas generators show near-zero progress and 92–100% stasis. A nine-annotator human study validates the ordering. The paper further shows that post-hoc frozen-readout optimization is gameable: the readout rises while an independent physical probe stays flat. As an alternative, it enforces monotonicity by construction in a disentangled attribute latent, validating this in a controlled renderer and on Stable-Diffusion semi-synthetic content, with elementary propositions in the supplement.","tokens_in":23986,"tokens_out":5894,"duration_ms":72428,"significance":"If the diagnosis holds, the paper contributes a reusable, null-robust evaluation protocol for irreversibility in video generation and a well-supported negative finding: current generators under-develop irreversible attributes rather than systematically reversing them. The paper is unusually careful in several ways: it null-tests its own proposed metrics (Tables I and II), runs real reference footage through the identical pipeline, uses bootstrap CIs, category-level paired tests, leave-one-category-out checks, and a human study, and it distinguishes between readout, probe, and ground truth. The guiding assumptions are stated and empirically checked in controlled settings, and the theoretical claims are elementary but explicit. The main weakness is that the central large-scale diagnosis relies on a CLIP directional readout whose attribute-specificity on natural video is not established against the confound of general dynamicness; this is load-bearing and fixable with additional controls, so the paper merits major revision rather than rejection.","major_comments":[{"comment":"The central diagnosis uses the per-category CLIP directional readout a(x)=<phi(x),u> as the measure of irreversible-attribute progress. This readout can in principle increase under any visual change with a component along u—lighting drift, camera motion, global color shifts—not only under the named attribute. Section III-f itself reports that generated clips make 2–3x less overall change than real reference footage, so the observed separation (ρ=+0.40 vs. ≈0; stasis 35% vs. 92–100%) could be partly or entirely a separation in general dynamicness rather than in attribute-specific development. The human validation in III-m uses a 0–4 rating of 'how much the process advances,' which can track the same dynamicness. No control separates the attribute-specific direction from generic directional change: e.g., a reversible attribute axis, an axis-orthogonal direction u_perp, motion-matched real/","section":"III-l and III-m (Table III)"},{"comment":"The readout-validity study in III-h is carried out on text-embedding-interpolated graded sequences (fixed latent, only the prompt attribute varies). This construct guarantees that the only systematic variation is the intended attribute and does not test discriminant validity against non-attribute sources of visual change (motion, camera, lighting, quality) in natural videos. Since the seven-model diagnosis is applied to real reference footage and generated clips with substantial motion differences, the protocol needs a demonstration that a(x) is insensitive to non-attribute change—for instance, a null experiment on reversed or frame-shuffled clips, or a comparison of a(x) with an independent probe on a sample of real videos. As written, the measurement study supports the readout's sensitivity to controlled attribute changes but not its specificity on the data to which the diagnosis is ap","section":"III-h"}],"minor_comments":[{"comment":"Several section references such as 'Sec. III-0d' and 'Sec. III-l' appear to be formatting artifacts from auto-numbering; please fix.","section":"Throughout"},{"comment":"The CogVideoX-2b row reports ρ=+0.16 and stasis 90% without a confidence interval, unlike the other rows. State how the 40 clips (5 processes x 8 seeds) are aggregated and whether the other models' clips are per-prompt independent.","section":"Table III"},{"comment":"Section III-a mentions 'human evaluation (a three-annotator study we run below)' but Section III-m describes nine annotators; reconcile the numbers.","section":"III-a and III-m"},{"comment":"The section title says 'adversarially gamed,' but the body qualifies this for in-loop guidance (perceptibly directional but partial, with identity preservation underpowered). Align the title or abstract wording with the qualified claim.","section":"IV-b and IV-c"},{"comment":"In Proposition 3, the constants m and κ are estimated through proxy probes; this limitation is stated in the supplement but should appear in the main text where the proposition is invoked.","section":"S2"}],"recommendation":"major_revision","confidential_remarks":"I would not reject on the confound issue: the paper's null-testing and scoping are exemplary, and the requested control experiments are feasible within the paper's existing framework. However, the central 'under-development of irreversible attributes' claim is currently entangled with general dynamicness in the seven-model study. The fit with IEEE TMM is acceptable, though the paper is more of a measurement-protocol-plus-mechanism contribution than a typical TMM application; the authors should make the protocol's attribute-specificity defense a priority in revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing worth knowing: this paper supplies a null-checked measurement protocol for irreversible dynamics in T2V, and a seven-model diagnosis that generators under-develop rather than reverse irreversible attributes. That diagnosis is much more carefully built than the typical eval paper — bootstrap CIs, category-level paired tests, leave-one-category-out, and a blind human study with inter-rater agreement. I buy the stasis finding; the attribute-specific reading is shakier.\n\nWhat's genuinely new: they show that per-clip normalized violation rates and variance-normalized reversal residuals are null-degenerate (they return ~0.5 and ~0.85 on pure noise), and that generic monotonicity scores reward stasis. Those are useful caveats for anyone building evaluation benchmarks. The gaming demonstration — post-hoc readout guidance closes the readout gap but not an independent physical probe — is convincing, and the proposition that readout-only objectives cannot certify the attribute is a clean theoretical statement.\n\nThe main soft spot is the readout. The seven-model separation uses per-category CLIP directional axes, and the paper itself reports generated clips move 2–3x less overall. So the real/generated gap in progress and stasis could be partly a dynamicness gap. The human validation uses a 0–4 'how much does the process advance' rating, which can also track general motion. There is no non-attribute control (lighting, camera, quality) to show the readout is sensitive to the attribute rather than to any directional change. The synthetic readout-validation study (rank corr ~0.84) helps, but it's generated content with controlled attribute interpolation, not the messy seven-model setting.\n\nOther soft spots are minor: Table II shows only null reads for the reversal residual, no positive control; no code or data release, which hampers reproducibility of the null simulations and human study; and the monotone-by-construction mechanism is validated only in synthetic and semi-synthetic settings — the authors are open about that, so it reads as a proof of concept, not a deployable fix.\n\nOverall, the central claim — under-development rather than reversal — survives the readout concern better than the stress-test note suggests. Even if the CLIP readout partly tracks dynamicness, the stasis result is robust because it aligns with the paper's independent finding that generated clips have 2–3x less net progress. What remains uncertain is exactly how much of the gap is attribute-specific. That's a fixable weakness, not a fatal one.\n\nThis is worth a serious referee slot. Send it to review with a request for a non-attribute motion control and a positive control for the reversal-residual null claim. I'd cite the protocol if I were working on video evaluation.","headline":"A null-checked protocol and seven-model diagnosis of under-development in T2V; the readout-validity gap is manageable but should be controlled before publication.","tokens_in":24439,"tokens_out":4402,"would_cite":true,"duration_ms":51231,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Text-to-video models under-develop irreversible attributes instead of reversing them, shown by a null-tested progress/stasis protocol.","keywords":["text-to-video generation","irreversible processes","temporal consistency","evaluation protocol","null baselines","under-development","disentangled representations","monotonicity"],"falsifier":"Take a clip of an object whose attribute never changes but whose camera moves steadily; if the directional text-based readout rises with the motion, the metric is not measuring the attribute, and the real-vs-generated separation would need to be re-scored.","tokens_in":23606,"feed_emoji":"⏳","tokens_out":9841,"duration_ms":116760,"temperature":0.7,"pith_summary":"Video generators are often treated as implicit world models, so whether they render irreversible changes—ice melting, paper charring, fruit rotting—is a basic fidelity question. The paper shows this question has been hard to answer: common metrics for 'reversal' score pure noise at chance and reward static clips. After null-testing candidate statistics, it builds a two-part protocol (attribute progress and stasis rate) and finds that seven text-to-video models barely develop irreversible attributes whereas real footage advances, a gap that nine human raters confirm. The reliable failure is under-development, not time reversal. A complementary result shows why the obvious fix—steering generation with a frozen attribute readout—is gameable, and why enforcing monotonicity inside a disentangled attribute latent can work instead.","feed_headline":"Seven video models stall; real time-lapse advances","feed_subtitle":"A null-tested measure finds generators barely develop rust, melt or rot; humans rate real clips 2.75 vs 0.99.","key_machinery":"Central machinery: a directional attribute readout a(x) = <phi(x), u>, where phi is a vision-language image encoder and u is a text-defined attribute direction; a progress–stasis protocol built from the rank correlation between time and the readout plus a stasis rate, both calibrated against null baselines; and, for the repair, a swap-consistency trained disentangled autoencoder D(z_a, z_c) whose scalar or vector attribute latent z_a is forced monotone by isotonic projection or softplus increment dynamics. The readout carries the diagnosis; the monotone latent carries the proposed construction; an independent probe (e.g. a color statistic) certifies that apparent changes are real.","core_discovery":"The paper's central claim is that current text-to-video generators fail irreversible processes by under-development, not by systematic time reversal, and that this can be established reliably only with a null-tested two-part protocol. Using a per-attribute directional readout (a vision-language similarity to prompts like 'rusted' minus 'clean') and running real reference footage through the identical pipeline, the authors report real clips advance (rho = +0.40, 35% static) while every one of seven generators shows near-zero progress and 92–100% stasis; nine annotators rate real clips 2.75 vs 0.99 on a 0–4 scale of how much the process advances. Naive reversal metrics—a per-clip violation rat","pith_inferences":["If the stasis finding generalizes, the right training fix is an explicit progress reward or a prompt-conditioned progress floor, not just an arrow-of-time constraint; a static clip already satisfies monotonicity.","The gaming result suggests a cheap safeguard for text-to-video evaluation: keep the score used to steer a model separate from the score used to judge it, and include a readout-independent physical or held-out probe.","The disentanglement bottleneck—leaky nuisance codes on real content—identifies a concrete research target: better attribute/nuisance separation on natural video could move the repair mechanism from synthetic renderers to real generators.","The null-degeneracy of normalized reversal metrics may apply beyond video generation, to any near-static time series where a monotone signal is measured with noise."],"forward_implications":["Any evaluation of irreversibility in text-to-video that reports a normalized violation rate or a generic embedding-monotonicity score should be re-run with the progress/stasis protocol; those metrics score pure noise and static clips favorably.","Unless training or architecture changes reward directional progress, new text-to-video models will likely continue to under-develop irreversible attributes; stasis is the dominant failure mode, not reversal.","Readout-guided editing or sampling that optimizes a frozen vision-language score cannot certify that the attribute was added; an independent probe is required, and even then the readout gap can close while the attribute stays flat.","Monotone-by-construction repair is only as good as the disentanglement: without a faithful attribute latent, a perfectly monotone latent can leave the visible attribute unchanged.","The protocol's progress and stasis numbers can serve as a reusable measurement instrument for future text-to-video models."],"supporting_citations":[{"why":"Supplies the real time-lapse reference set and six of the seven open-source generators used in the cross-model diagnosis.","marker":"[2]"},{"why":"Benchmarks object state change in text-to-video and motivates the need to measure irreversible directionality.","marker":"[1]"},{"why":"Provides the vision-language image and text encoders for the directional attribute readout used across the protocol and the guidance experiments.","marker":"[32]"},{"why":"Used for the five-process, eight-seed progress study that supports the under-development finding.","marker":"[35]"},{"why":"Latent diffusion renderer used to build nuisance-by-attribute grids for the semi-synthetic validation.","marker":"[34]"},{"why":"Adversarial diffusion distilled model used as the fast text-to-image renderer in the semi-synthetic pipeline.","marker":"[33]"}],"fun_headline_variants":["Video AI freezes time: real clips progress, models stall","7 video models show zero progress on rust, melt, rot","Irreversible physics stumps video generators: stasis 92–100%","Real time-lapses advance; AI video stalls at 0 progress","Under-development, not reversal: how video models fail physics"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The central diagnosis assumes the per-attribute measurement—comparing frames against text descriptions like 'rusted' versus 'clean'—tracks true attribute progress in both real and generated clips; if it partly responds to camera motion, lighting, or image quality, the reported gap between real and generated footage would shrink.","fun_headline_variants_meta":{"raw":{"variants":["Video AI freezes time: real clips progress, models stall","7 video models show zero progress on rust, melt, rot","Irreversible physics stumps video generators: stasis 92–100%","Real time-lapses advance; AI video stalls at 0 progress","Under-development, not reversal: how video models fail physics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":1081,"prompt_tokens":812,"completion_tokens":269,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":556,"completion_tokens_details":{"reasoning_tokens":179}},"tokens_in":556,"tokens_out":269,"duration_ms":3800,"temperature":1.0,"reasoning_tokens":179,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T00:31:23.488182+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a clip of an object whose attribute never changes but whose camera moves steadily; if the directional text-based readout rises with the motion, the metric is not measuring the attribute, and the real-vs-generated separation would need to be re-scored.","supporting_citations":[{"cited_title":"OSCBench: Benchmarking Object State Change in Text-to-Video Generation","cited_arxiv_id":"2603.11698","evidence_quote":"Benchmarks object state change in text-to-video and motivates the need to measure irreversible directionality."}],"review_version":1}