REVIEW 5 major objections 5 minor 112 references
Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance
T0 review · 5 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A self-supervised compression objective trained on 0.78 million unlabeled surgical frames produces an encoder that fine-tunes to improved results across surgical phase recognition, action triplet detection, segmentation, and polyp…
desk verdict Potentially useful large-scale surgical pretraining study undermined by a disconnect between its theoretical framing and its implemented MAE-style loss. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the compression encoder: each layer removes a constant dimension and applies a shrinkage step derived from the objective $\max_\theta\,-H(Z)+H(I)$, which the paper connects to Kolmogorov complexity through the asymptotic identity $E[\frac{1}{n}K(Z^n)]\to H(Z)$. The gradient update uses an SVD-based form $Z_{i+\frac12}=Z_{i+1}-\beta(S_{i+1}V_{i+1}^{-1}D_{i+1})$ with a residual-like projection bypass, so each layer is meant to solve a concave entropy-maximization subproblem. The decoder inverts this with a temperature-dependent Langevin sampler, where the temperature is estimated from a conditional slice of the hidden state. In the actual pretraining, however, the stated loss is mean absolute error between reconstructed and original patches, run as a standard masked autoencoder pipeline.
What would settle it
Train the same architecture with the same data, masking, and training schedule but replace the decoder and losses with a plain MAE decoder and L2 or MAE reconstruction; if phase, action-triplet, segmentation, and polyp metrics are statistically indistinguishable from C2E's, the compression and entropy machinery is not carrying the reported gains. Conversely, the paper's claim would be supported by showing that removing the entropy-maximizing Langevin decoder or the dimension-reduction projections measurably degrades downstream accuracy.
Extended reading notes
Core claim
On the paper's own terms, C2E establishes that maximizing the entropy difference between input and a dimension-reduced hidden state, equivalently minimizing Kolmogorov complexity under the identification of complexity with entropy, yields an encoder whose latent representations are compact and disentangled by surgery type, phase, and instrument. The reconstruction side uses a decoder derived from the closed-form entropy-maximizing distribution $p(z)\propto e^{-E(z)/E[E(z)]}$ sampled via Langevin dynamics, intended to recover perceptually fine details such as tool textures and organ boundaries. The paper reports, with the encoder used as a backbone, 92.5% accuracy on Cholec80 phase recognition, 41.8 mAP on action triplet recognition, 0.74 IoU on CholecSeg8k segmentation, and 94.0% accuracy on polyp diagnosis, and substantially higher few-shot phase accuracy than a vision transformer pretrained on endoscopic images.
Load-bearing premise
The pretraining that produced the reported results is the same network described by the compression and entropy-maximization objectives, rather than a standard masked-autoencoder reconstruction loss with the theoretical machinery playing no functional role.
Editorial extensions
If this is right
- If C2E's compression objective is what drives its representations, unlabeled video libraries become the main resource for building surgical foundation models rather than labeled benchmarks.
- Fine-tuning C2E on as few as two or eight videos yields phase recognition gains of roughly 8 to 26 points over a ViT baseline on colorectal procedure types, suggesting new procedures can be modeled with minimal annotation.
- Because one pretrained encoder helps classification, triplet detection, segmentation, and diagnosis, downstream systems can share a single backbone instead of task-specific pretraining.
- The claimed disentanglement of surgery type, phase, and instrument in the latent space could enable interpretable monitoring, with attention heads tracking local details and their aggregation composing global context.
Reading between the lines
- The reported comparisons mostly keep the downstream decoder fixed while swapping the backbone, so the incremental gains may come from the larger and more diverse pretraining corpus rather than from the compression theory; a controlled experiment pretraining plain MAE on the identical corpus would separate the two.
- If the effective objective is standard MAE, C2E can be read as evidence that scaling up diverse surgical videos with a masked-autoencoder objective is sufficient, with the Kolmogorov-complexity framing serving as an interpretation rather than a mechanism.
- A testable extension is to evaluate C2E on other fine-grained medical video domains, such as endoscopy or retinal surgery, where local texture is decisive; the paper's own logic predicts the largest gains where global and local features must be balanced.
- The temperature-conditional decoder suggests a route to per-image uncertainty estimates in downstream predictions, though the paper does not attempt this.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Compress-to-Explore (C2E), a self-supervised pretraining framework for surgical video understanding. The method claims a compression encoder that maximizes a Kolmogorov-complexity difference objective and an exploration decoder based on entropy-maximizing Langevin sampling, pretrained on 0.78M unlabeled surgical images from 2122 procedures. The pretrained encoder is then fine-tuned on phase recognition, action-triplet classification, segmentation, polyp diagnosis, and few-shot phase recognition. The central claims are that C2E produces compact, disentangled representations and improves downstream accuracy relative to ViT/MAE-style baselines, thereby serving as a surgical visual foundation model.
Significance. If the empirical results were reproducible and the mechanism were actually implemented, the scale of pretraining (0.78M frames, 2122 surgeries) and the breadth of downstream tasks would be a useful contribution to surgical computer vision. The paper also includes few-shot evaluations and representation visualizations, which are appropriate for a foundation-model claim. However, the theoretical core supporting the C2E mechanism is not sound, the implemented loss appears to be standard MAE reconstruction, and the empirical comparisons do not consistently show the claimed improvements. The paper does not release code or the private MGH dataset, so the reported gains cannot currently be attributed to the proposed mechanism rather than to data scale or an architectural variant of MAE. The significance as stated, namely a new compression- and entropy-based self-supervised principle, is therefore not established.
major comments (5)
- [Section II-D] The paper explicitly states that 'the loss function minimizes the mean average error of the predicted images and the original images, making the overall training run as a standard MAE [17] pipeline.' None of the terms in Eq. 3, such as -1/2 ln|Σ_Z| and the dimension-change term, appears in this loss, and Eq. 10's energy, temperature kT(h), and Langevin noise are likewise absent from any weight update. Consequently, the trained encoder is not the one whose compression and entropy-maximization properties are proved in Section V. This is a load-bearing inconsistency: the central claim that C2E 'leverages Kolmogorov complexity' and 'uses entropy-maximizing decoders' is disconnected from the implemented training objective.
- [Section V, Theorems 1 and 3] Theorem 1 states that the entropy H(X) is the Kolmogorov complexity K(X), but the proof establishes only the asymptotic relation H(X) ≤ E[(1/n)K(X)] ≤ H(X) + ... for finite alphabets, i.e., that the expected normalized Kolmogorov complexity converges to H(X). This does not imply equality of H(X) and K(X) for individual embeddings. In addition, Theorem 3 refers to H(Z) as a 'higher bound' of K(X), while Eq. 13 places H(X) as a lower bound of the expected complexity; the inequality is used in the wrong direction. These errors undermine the theoretical foundation of the compression encoder.
- [Section II-B, Eq. (4), and Section V, Theorems 2 and 5] The algebraic derivation of Eq. 4 is not valid as written. From Z = SVD, the simplification Z(Z^T Z)^{-1} = S V^{-1} D requires orthogonality and invertibility assumptions that are not stated and are not satisfied by the rectangular matrix Z ∈ R^{N,C}; the computation also drops D^{-1} and misplaces the transpose structure. Separately, Theorems 2 and 5 invoke the concavity of ln|A| on positive definite Hermitian matrices and assert that 'Z is a positive definite Hermitian matrix', but the paper's own notation defines Z as an N×C matrix of hidden states. The matrix-analysis theorems therefore do not apply as stated, leaving Eqs. 3-5 without a valid justification.
- [Section II-C, Eqs. (6)-(8)] The maximum-entropy solution is misidentified in Eqs. 6-8. Under a constraint on E[E(z)], the Lagrange multiplier λ is the inverse temperature, and the solution is p(z) ∝ exp(-λ E(z)); it is not p(z) ∝ exp(-E(z)/E[E(z)]). Writing E[E(z)] = kT(h) in Eq. 8 conflates the expectation of the energy with the temperature parameter, yet this identification is used directly in the Langevin sampling of Eq. 10. The exploration decoder's theoretical grounding is therefore also unsupported.
- [Section III, Tables II-VI] The reported results do not consistently support the claim of improved performance. In Table II, C2E (92.5±6.9) is statistically indistinguishable from SurgFormer (92.4±6.4), and its precision is lower than several baselines. In Table III, C2E's mAPit (42.3±0.9) is below MT4MTL's (43.1±2.0), even though the text claims better average precision across the action-triplet metrics. More importantly, no ablation controls for the 531,565 private MGH images or for the specific architecture changes against a standard MAE trained on the same data; the only ViT comparison in Table VI uses a previously published model rather than a same-data, same-compute MAE baseline. Without such controls, the reported gains cannot be attributed to the C2E mechanism.
minor comments (5)
- [Throughout] There are numerous typos, including 'Massachusset' in the author affiliation, 'adoped' in Section III-D, 'few-short learning' in the contributions list and Section III-F, and 'Precession' instead of 'Precision' in the Table II header.
- [Table VII] The entry '58.8 ±11.' is incomplete; the standard deviation is missing its decimal digits.
- [Section III-F and Figure 3] The naming is inconsistent: 'VIT' is used in Figure 3 and Table VI while 'ViT' is used elsewhere; please unify.
- [Section II, notation] The symbol E is used both for the energy function and for the expectation operator, e.g., E[E(z)] in Eqs. 6-8 and Eq. 10; this is confusing and should be disambiguated with different symbols.
- [Section II-A] The claim that Z0 is a 'disentangled and sparse normal distribution' is not supported by any quantitative evaluation in the paper; the t-SNE visualizations in Section III-F are not a substitute for a measured sparsity or disentanglement metric.
Circularity Check
No significant circularity: downstream results are external-benchmark evaluations with no labels in pretraining; the theory–implementation gap is a validity issue, not a circular reduction.
full rationale
The derivation chain is not circular. Pretraining in Section II-D is described as a standard MAE reconstruction pipeline over 0.78M unlabeled frames, and all reported downstream numbers (Cholec80 phase recognition, CholecT45 action triplets, CholecSeg8k segmentation, PolypDiag diagnosis, and few-shot HeiCo evaluations) are measured against external public benchmarks and external baseline methods, with no downstream label used during pretraining and with claimed removal of overlap images to avoid leakage. The compression and entropy-maximization equations (Eq. 3 and Eq. 10) are presented as motivation for the encoder/decoder architecture, and even if the paper's own statement that training 'runs as a standard MAE pipeline' means those information-theoretic objectives are not separately optimized, that is an implementation/validity gap rather than a circular reduction of the predicted result to a fitted input. The mathematical flaws in Theorem 1 (identifying expected Kolmogorov complexity with entropy despite only an asymptotic inequality) and Theorem 3 (calling H(Z) a 'higher bound' when Eq. 13 places it on the lower side) are correctness concerns, not circularity. Self-citations to [48], [61], [62], [66], and [111] appear only as related work or illustrative context and are not load-bearing for the empirical claims. No parameter fitted to a subset of downstream data is renamed as a prediction, and no uniqueness or forced-choice conclusion depends on an author-cited prior theorem. The paper is therefore best described as self-contained against external benchmarks, with the principal risks being theoretical overclaim and disconnect between stated mechanism and implemented loss, not circularity.
Assumptions & free parameters
free parameters (3)
- conditional ratio beta
- Langevin step size epsilon
- gradient step size beta in Eq. 4
assumptions (5)
- standard math Kolmogorov complexity of image embeddings asymptotically approaches entropy (Cover and Thomas Thm 14.3.1)
- domain assumption The covariance of embeddings is approximated by Sigma = (1/m) Z Z^T after zero-mean normalization
- ad hoc to paper Dimension-change term (1/2) Delta M can be dropped because dimension changes are constant
- ad hoc to paper Assumption 1 in Theorem 6: the distribution of particles z is homogeneous across all subspaces
- ad hoc to paper Z is a positive definite Hermitian matrix
Cite this review
Pith. "Pith review of Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance." pith.science (2026). https://pith.science/paper/LRVYWDCU
@misc{pith2026250601980,
author = {Pith},
title = {Pith review of: Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance},
year = {2026},
howpublished = {\url{https://pith.science/paper/LRVYWDCU}},
note = {Machine review of arXiv:2506.01980}
}
read the original abstract
Real-time video understanding is critical to guide procedures in minimally invasive surgery (MIS). However, supervised learning approaches require large, annotated datasets that are scarce due to annotation efforts that are prohibitive, e.g., in medical fields. Although self-supervision methods can address such limitations, current self-supervised methods often fail to capture structural and physical information in a form that generalizes across tasks. We propose Compress-to-Explore (C2E), a novel self-supervised framework that leverages Kolmogorov complexity to learn compact, informative representations from surgical videos. C2E uses entropy-maximizing decoders to compress images while preserving clinically relevant details, improving encoder performance without labeled data. Trained on large-scale unlabeled surgical datasets, C2E demonstrates strong generalization across a variety of surgical ML tasks, such as workflow classification, tool-tissue interaction classification, segmentation, and diagnosis tasks, providing improved performance as a surgical visual foundation model. As we further show in the paper, the model's internal compact representation better disentangles features from different structural parts of images. The resulting performance improvements highlight the yet untapped potential of self-supervised learning to enhance surgical AI and improve outcomes in MIS.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[17]
Masked autoencoders are scalable vision learners,
K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pp. 16000–16009, 2022
2022
-
[1]
Automated 3d liver segmentation from hepatobiliary phase mri for enhanced preoperative planning,
N. Oh, J.-H. Kim, J. Rhu, W. K. Jeong, G.-s. Choi, J. M. Kim, and J.-W. Joh, “Automated 3d liver segmentation from hepatobiliary phase mri for enhanced preoperative planning,” Scientific Reports , vol. 13, no. 1, p. 17605, 2023
2023
-
[2]
Ai-based chest ct semantic segmentation algorithm enables semi-automated lung cancer surgery planning by recognizing anatomical variants of pulmonary vessels,
X. Chen, H. Xu, Q. Qi, C. Sun, J. Jin, H. Zhao, X. Wang, W. Weng, S. Wang, X. Sui, et al. , “Ai-based chest ct semantic segmentation algorithm enables semi-automated lung cancer surgery planning by recognizing anatomical variants of pulmonary vessels,” Frontiers in Oncology, vol. 12, p. 1021084, 2022
2022
-
[3]
Image-guided simulation of heterogeneous tissue deforma- tion for augmented reality during hepatic surgery,
N. Haouchine, J. Dequidt, I. Peterl ´ık, E. Kerrien, M. Berger, and S. Cotin, “Image-guided simulation of heterogeneous tissue deforma- tion for augmented reality during hepatic surgery,” 2013
2013
-
[4]
Computer vision in surgery: from potential to clinical value,
P. Mascagni, D. Alapatt, L. Sestini, M. S. Altieri, A. Madani, Y . Watan- abe, A. Alseidi, J. A. Redan, S. Alfieri, G. Costamagna, et al. , “Computer vision in surgery: from potential to clinical value,” npj Digital Medicine, vol. 5, no. 1, p. 163, 2022
2022
-
[5]
Laparoscopic image-based critical action recognition and anticipation with explainable features,
J. Zhang, S. Zhou, Y . Wang, S. Shi, C. Wan, H. Zhao, X. Cai, and H. Ding, “Laparoscopic image-based critical action recognition and anticipation with explainable features,” IEEE J. Biomed. Health Inform., vol. 27, pp. 5393–5404, Nov. 2023
2023
-
[6]
Current and future applications of artificial intelligence in surgery: implications for clinical practice and research,
M. X. Morris, D. Fiocco, T. Caneva, P. Yiapanis, and D. P. Orgill, “Current and future applications of artificial intelligence in surgery: implications for clinical practice and research,” Front. Surg., vol. 11, p. 1393898, May 2024
2024
-
[7]
Artificial intelligence in surgery: the future is now,
H. Ashrafiana, “Artificial intelligence in surgery: the future is now,” Eur Surg Res , vol. 65, pp. 22–39, 2024
2024
Show all 112 references
-
[8]
Artificial in- telligence in improving the outcome of surgical treatment in colorectal cancer,
M. F. Avram, D. C. Laz ˘ar, M. I. Maris ¸, and S. Olariu, “Artificial in- telligence in improving the outcome of surgical treatment in colorectal cancer,” Front. Oncol., vol. 13, p. 1116761, Jan. 2023
2023
-
[9]
Shalev-Shwartz and S
S. Shalev-Shwartz and S. Ben-David, Understanding machine learning: From theory to algorithms . Cambridge university press, 2014
2014
-
[10]
SAGES consensus recommendations on an annotation framework for surgical video,
O. R. Meireles, G. Rosman, M. S. Altieri, L. Carin, G. Hager, A. Madani, N. Padoy, C. M. Pugh, P. Sylla, T. M. Ward, D. A. Hashimoto, and SAGES Video Annotation for AI Working Groups, “SAGES consensus recommendations on an annotation framework for surgical video,” Surg. Endosc...
2021
-
[11]
J. A. Eckhoff, G. Rosman, M. S. Altieri, S. Speidel, D. Stoyanov, M. Anvari, L. Meier-Hein, K. M ¨arz, P. Jannin, C. Pugh, M. Wagner, E. Witkowski, P. Shaw, A. Madani, Y . Ban, T. Ward, F. Filicori, N. Padoy, M. Talamini, and O. R. Meireles, “SAGES consensus recommendations on...
2023
-
[12]
What does classifying more than 10,000 image categories tell us?,
J. Deng, A. C. Berg, K. Li, and L. Fei-Fei, “What does classifying more than 10,000 image categories tell us?,” inComputer Vision–ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece, September 5-11, 2010, Proceedings, Part V 11 , pp. 71– 84, Springer, 2010
2010
-
[13]
MT4MTL- KD: A multi-teacher knowledge distillation framework for triplet recognition,
S. Gui, Z. Wang, J. Chen, X. Zhou, C. Zhang, and Y . Cao, “MT4MTL- KD: A multi-teacher knowledge distillation framework for triplet recognition,” IEEE Trans. Med. Imaging, vol. 43, pp. 1628–1639, Apr. 2024
2024
-
[14]
Foundation model for endoscopy video analysis via large-scale self-supervised pre-train,
Z. Wang, C. Liu, S. Zhang, and Q. Dou, “Foundation model for endoscopy video analysis via large-scale self-supervised pre-train,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2023, pp. 101–111, Springer Nature Switzerland, 2023
2023
-
[15]
The effect of image resolution on deep learning in radiography,
C. F. Sabottke and B. M. Spieler, “The effect of image resolution on deep learning in radiography,” Radiol. Artif. Intell., vol. 2, p. e190015, Jan. 2020
2020
-
[16]
Ultra-high resolution, multi-scale, context- aware approach for detection of small cancers on mammography,
K. Rangarajan, A. Gupta, S. Dasgupta, U. Marri, A. K. Gupta, S. Hari, S. Banerjee, and C. Arora, “Ultra-high resolution, multi-scale, context- aware approach for detection of small cancers on mammography,” Sci. Rep., vol. 12, p. 11622, July 2022
2022
-
[18]
Exploring the effect of dataset diversity in self-supervised learning for surgical computer vision,
T. J. Jaspers, R. L. de Jong, Y . Al Khalil, T. Zeelenberg, C. H. Kusters, Y . Li, R. C. van Jaarsveld, F. H. Bakker, J. P. Ruurda, W. M. Brinkman, et al. , “Exploring the effect of dataset diversity in self-supervised learning for surgical computer vision,” in MICCAI Workshop...
2024
-
[19]
Endovit: pretraining vision transformers on a large collection of endoscopic images,
D. Bati ´c, F. Holm, E. ¨Ozsoy, T. Czempiel, and N. Navab, “Endovit: pretraining vision transformers on a large collection of endoscopic images,” International Journal of Computer Assisted Radiology and Surgery, vol. 19, no. 6, pp. 1085–1091, 2024
2024
-
[20]
High-resolution image synthesis with latent diffusion models,
R. Rombach, A. Blattmann, D. Lorenz, and others, “High-resolution image synthesis with latent diffusion models,” Proceedings of the , 2022
2022
-
[21]
On the opportunities and risks of foundation models,
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al. , “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258, 2021
2021 arXiv
-
[22]
Exploring simple siamese representation learn- ing,
X. Chen and K. He, “Exploring simple siamese representation learn- ing,” arXiv [cs.CV], pp. 15750–15758, Nov. 2020
2020
-
[23]
Self-supervised learning for endoscopic video analysis,
R. Hirsch, M. Caron, R. Cohen, A. Livne, R. Shapiro, T. Golany, R. Goldenberg, D. Freedman, and E. Rivlin, “Self-supervised learning for endoscopic video analysis,” arXiv [cs.CV], Aug. 2023
2023
-
[24]
Pseudo- label guided cross-video pixel contrast for robotic surgical scene seg- mentation with limited annotations,
Y . Yu, Z. Zhao, Y . Jin, G. Chen, Q. Dou, and P.-A. Heng, “Pseudo- label guided cross-video pixel contrast for robotic surgical scene seg- mentation with limited annotations,” in 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 10857– 1086...
2022
-
[25]
Less is more: Sur- gical phase recognition with less annotations through self-supervised pre-training of cnn-lstm networks,
G. Yengera, D. Mutter, J. Marescaux, and N. Padoy, “Less is more: Sur- gical phase recognition with less annotations through self-supervised pre-training of cnn-lstm networks,” arXiv preprint arXiv:1805.08569 , 2018
2018 arXiv
-
[26]
“train one, classify one, teach one
D. Neimark, O. Bar, M. Zohar, G. D. Hager, and D. Asselmann, ““train one, classify one, teach one”-cross-surgery transfer learning for surgical step recognition,” in Medical Imaging with Deep Learning , pp. 532– 544, PMLR, 2021
2021
-
[27]
Dissecting self- supervised learning methods for surgical computer vision,
S. Ramesh, V . Srivastav, D. Alapatt, T. Yu, A. Murali, L. Sestini, C. I. Nwoye, I. Hamoud, S. Sharma, A. Fleurentin, et al., “Dissecting self- supervised learning methods for surgical computer vision,” Medical Image Analysis, vol. 88, p. 102844, 2023
2023
-
[28]
Self-supervised learning from images with a joint-embedding predictive architecture,
M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, Y . LeCun, and N. Ballas, “Self-supervised learning from images with a joint-embedding predictive architecture,” arXiv [cs.CV], pp. 15619– 15629, Jan. 2023
2023
-
[29]
VICRegL: Self-supervised learning of local visual features,
A. Bardes, J. Ponce, and Y . LeCun, “VICRegL: Self-supervised learning of local visual features,” in Advances in Neural Information Processing Systems (S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, eds.), vol. 35, pp. 8799–8810, Curran Associates, Inc., 2022
2022
-
[30]
Denoising diffusion probabilistic models,
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” arXiv [cs.LG], pp. 6840–6851, June 2020
2020
-
[31]
Diffusion models for medical image analysis: A comprehensive survey,
A. Kazerouni, E. K. Aghdam, M. Heidari, R. Azad, M. Fayyaz, I. Hacihaliloglu, and D. Merhof, “Diffusion models for medical image analysis: A comprehensive survey,” arXiv [eess.IV], Nov. 2022
2022
-
[32]
A fast karhunen-loeve transform for a class of random processes,
A. K. Jain, “A fast karhunen-loeve transform for a class of random processes,” IEEE Transactions on Communications , vol. 24, no. 9, pp. 1023–1029, 1976
1976
-
[33]
Segmentation of multivariate mixed data via lossy data coding and compression,
Y . Ma, H. Derksen, W. Hong, and J. Wright, “Segmentation of multivariate mixed data via lossy data coding and compression,” IEEE transactions on pattern analysis and machine intelligence , vol. 29, no. 9, pp. 1546–1562, 2007
2007
-
[34]
Provable bounds for learning some deep representations,
S. Arora, A. Bhaskara, R. Ge, and T. Ma, “Provable bounds for learning some deep representations,” in International conference on machine learning, pp. 584–592, PMLR, 2014
2014
-
[35]
White-box transformers via sparse rate reduction,
Y . Yu, S. Buchanan, D. Pai, T. Chu, Z. Wu, S. Tong, B. Haeffele, and Y . Ma, “White-box transformers via sparse rate reduction,”Advances in Neural Information Processing Systems, vol. 36, pp. 9422–9457, 2023
2023
-
[36]
Deep learning and the information bottleneck principle,
N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in 2015 ieee information theory workshop (itw) , pp. 1–5, Ieee, 2015. 10
2015
-
[37]
Direct validation of the information bottleneck principle for deep nets,
A. Elad, D. Haviv, Y . Blau, and T. Michaeli, “Direct validation of the information bottleneck principle for deep nets,” 2019
2019
-
[38]
The information bottleneck problem and its applications in machine learning,
Z. Goldfeld and Y . Polyanskiy, “The information bottleneck problem and its applications in machine learning,” 2020
2020
-
[39]
On neural networks fitting, compression, and generalization behavior via information-bottleneck- like approaches,
Z. Lyu, G. Aminian, and M. Rodrigues, “On neural networks fitting, compression, and generalization behavior via information-bottleneck- like approaches,” 2023
2023
-
[40]
Neural manifold clustering and embedding,
Z. Li, Y . Chen, Y . LeCun, and F. T. Sommer, “Neural manifold clustering and embedding,” arXiv preprint arXiv:2201.10000 , 2022
2022 arXiv
-
[41]
Self-supervised learning via maximum entropy coding,
X. Liu, Z. Wang, Y .-L. Li, and S. Wang, “Self-supervised learning via maximum entropy coding,” Advances in neural information processing systems, vol. 35, pp. 34091–34105, 2022
2022
-
[42]
Segmentation of multivariate mixed data via lossy data coding and compression,
Y . Ma, H. Derksen, and W. Hong, “Segmentation of multivariate mixed data via lossy data coding and compression,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 29, pp. 1546–1562, Sept. 2007
2007
-
[43]
Fast: Efficient action tokenization for vision-language-action models,
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine, “Fast: Efficient action tokenization for vision-language-action models,” arXiv preprint arXiv:2501.09747 , 2025
2025 arXiv
-
[44]
Human-like systematic generalization through a meta-learning neural network,
B. M. Lake and M. Baroni, “Human-like systematic generalization through a meta-learning neural network,” Nature, vol. 623, no. 7985, pp. 115–121, 2023
2023
-
[45]
Unsuper- vised learning of compositional energy concepts,
Y . Du, S. Li, Y . Sharma, J. Tenenbaum, and I. Mordatch, “Unsuper- vised learning of compositional energy concepts,” Advances in Neural Information Processing Systems , vol. 34, pp. 15608–15620, 2021
2021
-
[46]
Energy-based models are zero-shot planners for compositional scene rearrangement,
N. Gkanatsios, A. Jain, Z. Xian, Y . Zhang, C. Atkeson, and K. Fragki- adaki, “Energy-based models are zero-shot planners for compositional scene rearrangement,” arXiv preprint arXiv:2304.14391 , 2023
2023 arXiv
-
[47]
Compositional visual generation with energy based models,
Y . Du, S. Li, and I. Mordatch, “Compositional visual generation with energy based models,” Advances in Neural Information Processing Systems, vol. 33, pp. 6637–6647, 2020
2020
-
[48]
Artificial intelligence in surgery: Promises and perils,
D. A. Hashimoto, G. Rosman, D. Rus, and O. R. Meireles, “Artificial intelligence in surgery: Promises and perils,” Ann. Surg. , vol. 268, pp. 70–76, July 2018
2018
-
[49]
Superpixel-based structure classification for laparoscopic surgery,
S. Bodenstedt, J. G ¨ortler, M. Wagner, H. Kenngott, B. M ¨uller-Stich, R. Dillmann, and S. Speidel, “Superpixel-based structure classification for laparoscopic surgery,” 2016
2016
-
[50]
ToolNet: Holistically-nested real-time segmentation of robotic surgical tools,
L. C. Garc ´ıa-Peraza-Herrera, W. Li, L. Fidon, C. Gruijthuijsen, A. Devreker, G. Attilakos, J. Deprest, E. V . Poorten, D. Stoyanov, T. Vercauteren, and S. Ourselin, “ToolNet: Holistically-nested real-time segmentation of robotic surgical tools,” in2017 IEEE/RSJ International...
2017
-
[51]
2018 robotic scene segmentation challenge,
M. Allan, S. Kondo, S. Bodenstedt, S. Leger, R. Kadkhodamoham- madi, I. Luengo, F. Fuentes-Hurtado, E. Flouty, A. K. Mohammed, M. Pedersen, A. Kori, A. Varghese, G. Krishnamurthi, D. Rauber, R. Mendel, C. Palm, S. Bano, G. Saibro, C. Shih, H. Chiang, J. Zhuang, J. Yang, V . Ig...
2018 arXiv
-
[52]
Cholecseg8k: a semantic segmentation dataset for laparoscopic cholecystectomy based on cholec80,
W.-Y . Hong, C.-L. Kao, Y .-H. Kuo, J.-R. Wang, W.-L. Chang, and C.-S. Shih, “Cholecseg8k: a semantic segmentation dataset for laparoscopic cholecystectomy based on cholec80,” arXiv preprint arXiv:2012.12453, 2020
2012 arXiv
-
[53]
Detection and localization of robotic tools in robot-assisted surgery videos using deep neural net- works for region proposal and detection,
D. Sarikaya, J. J. Corso, and K. A. Guru, “Detection and localization of robotic tools in robot-assisted surgery videos using deep neural net- works for region proposal and detection,” IEEE Trans. Med. Imaging , vol. 36, pp. 1542–1549, July 2017
2017
-
[54]
Tool detection and operative skill assessment in surgical videos using region-based convolutional neural networks,
A. Jin, S. Yeung, J. Jopling, J. Krause, D. Azagury, A. Milstein, and L. Fei-Fei, “Tool detection and operative skill assessment in surgical videos using region-based convolutional neural networks,” in 2018 IEEE Winter Conference on Applications of Computer Vision (WACV) , pp....
2018
-
[55]
LapFormer: surgical tool detection in laparoscopic surgical video using transformer architecture,
S. Kondo, “LapFormer: surgical tool detection in laparoscopic surgical video using transformer architecture,” Computer Methods in Biome- chanics and Biomedical Engineering: Imaging & Visualization , vol. 9, pp. 302–307, May 2021
2021
-
[56]
Vision-based and marker-less surgical tool detection and tracking: a review of the literature,
D. Bouget, M. Allan, D. Stoyanov, and P. Jannin, “Vision-based and marker-less surgical tool detection and tracking: a review of the literature,” Medical image analysis , vol. 35, pp. 633–654, 2017
2017
-
[57]
Tracking of in- struments in minimally invasive surgery for surgical skill analysis,
S. Speidel, M. Delles, C. Gutt, and R. Dillmann, “Tracking of in- struments in minimally invasive surgery for surgical skill analysis,” in Medical Imaging and Augmented Reality: Third International Work- shop, Shanghai, China, August 17-18, 2006 Proceedings 3 , pp. 148– 155, S...
2006
-
[58]
Tracking-by- detection of surgical instruments in minimally invasive surgery via the convolutional neural network deep learning-based method,
Z. Zhao, S. V oros, Y . Weng, F. Chang, and R. Li, “Tracking-by- detection of surgical instruments in minimally invasive surgery via the convolutional neural network deep learning-based method,” Computer Assisted Surgery, vol. 22, no. sup1, pp. 26–35, 2017
2017
-
[59]
Statistical modeling and recognition of surgical workflow,
N. Padoy, T. Blum, S.-A. Ahmadi, H. Feussner, M.-O. Berger, and N. Navab, “Statistical modeling and recognition of surgical workflow,” Medical image analysis , vol. 16, no. 3, pp. 632–641, 2012
2012
-
[60]
Predicting surgical phases using cnn-narx neural network,
N. A. Jalal, T. A. Alshirbaji, and K. M ¨oller, “Predicting surgical phases using cnn-narx neural network,” Current Directions in Biomedical Engineering, vol. 5, no. 1, pp. 405–407, 2019
2019
-
[61]
Supr-gan: Surgical prediction gan for event anticipation in laparoscopic and robotic surgery,
Y . Ban, G. Rosman, J. A. Eckhoff, T. M. Ward, D. A. Hashimoto, T. Kondo, H. Iwaki, O. R. Meireles, and D. Rus, “Supr-gan: Surgical prediction gan for event anticipation in laparoscopic and robotic surgery,” IEEE Robotics and Automation Letters , vol. 7, no. 2, pp. 5741–5748, 2022
2022
-
[62]
Hypergraph-transformer (hgt) for interactive event prediction in la- paroscopic and robotic surgery,
L. Yin, Y . Ban, J. Eckhoff, O. Meireles, D. Rus, and G. Rosman, “Hypergraph-transformer (hgt) for interactive event prediction in la- paroscopic and robotic surgery,” arXiv preprint arXiv:2402.01974 , 2024
2024 arXiv
-
[63]
Surgical SAM 2: Real- time segment anything in surgical video by efficient frame pruning,
H. Liu, E. Zhang, J. Wu, M. Hong, and Y . Jin, “Surgical SAM 2: Real- time segment anything in surgical video by efficient frame pruning,” arXiv preprint arXiv:2408.07931 , 2024
2024 arXiv
-
[64]
T. M. Cover and J. A. Thomas, Elements of Information Theory 2nd Edition (Wiley Series in Telecommunications and Signal Processing) . Wiley-Interscience, 2 ed., July 2006
2006
-
[65]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , pp. 770–778, 2016
2016
-
[66]
Fast regularization of matrix-valued images,
G. Rosman, Y . Wang, X.-C. Tai, R. Kimmel, and A. M. Bruckstein, “Fast regularization of matrix-valued images,” in Efficient Algorithms for Global Optimization Methods in Computer Vision: International Dagstuhl Seminar, Dagstuhl Castle, Germany, November 20-25, 2011, Revised S...
2011
-
[67]
Information theory and statistical mechanics,
E. T. Jaynes, “Information theory and statistical mechanics,” Physical review, 1957
1957
-
[68]
Imagededup
T. Jain, C. Lennan, Z. John, and D. Tran, “Imagededup.” https: //github.com/idealo/imagededup, 2019
2019
-
[69]
Exploring CLIP for assessing the look and feel of images,
J. Wang, K. C. Chan, and C. C. Loy, “Exploring CLIP for assessing the look and feel of images,” in AAAI, pp. 2555–2563, 2023
2023
-
[70]
EndoNet: A deep architecture for recognition tasks on laparoscopic videos,
A. P. Twinanda, S. Shehata, D. Mutter, J. Marescaux, M. de Mathelin, and N. Padoy, “EndoNet: A deep architecture for recognition tasks on laparoscopic videos,” IEEE Trans. Med. Imaging , vol. 36, pp. 86–97, Jan. 2017
2017
-
[71]
hsdb-instrument: Instrument localization database for laparoscopic and robotic surgeries,
J. Yoon, J. Lee, S. Heo, H. Yu, J. Lim, C. H. Song, S. Hong, S. Hong, B. Park, S. Park, et al. , “hsdb-instrument: Instrument localization database for laparoscopic and robotic surgeries,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th Internat...
2021
-
[72]
Heidelberg colorectal data set for surgical data science in the sensor operating room,
L. Maier-Hein, M. Wagner, T. Ross, A. Reinke, S. Bodenstedt, P. M. Full, H. Hempe, D. Mindroc-Filimon, P. Scholz, T. N. Tran, et al. , “Heidelberg colorectal data set for surgical data science in the sensor operating room,” Scientific data, vol. 8, no. 1, p. 101, 2021
2021
-
[73]
The dresden surgical anatomy dataset for abdominal organ segmentation in surgical data science,
M. Carstens, F. M. Rinner, S. Bodenstedt, A. C. Jenke, J. Weitz, M. Distler, S. Speidel, and F. R. Kolbinger, “The dresden surgical anatomy dataset for abdominal organ segmentation in surgical data science,” Scientific Data, vol. 10, no. 1, pp. 1–8, 2023
2023
-
[74]
Lapgyn4: a dataset for 4 automatic content analysis problems in the domain of laparoscopic gynecol- ogy,
A. Leibetseder, S. Petscharnig, M. J. Primus, S. Kletz, B. M ¨unzer, K. Schoeffmann, and J. Keckstein, “Lapgyn4: a dataset for 4 automatic content analysis problems in the domain of laparoscopic gynecol- ogy,” in Proceedings of the 9th ACM multimedia systems conference , pp. 3...
2018
-
[75]
Video retrieval in laparoscopic video recordings with dynamic content descriptors,
K. Schoeffmann, H. Husslein, S. Kletz, S. Petscharnig, B. Muenzer, and C. Beecks, “Video retrieval in laparoscopic video recordings with dynamic content descriptors,” Multimedia Tools and Applications, vol. 77, pp. 16813–16832, 2018
2018
-
[76]
Glenda: gynecologic laparoscopy endometriosis dataset,
A. Leibetseder, S. Kletz, K. Schoeffmann, S. Keckstein, and J. Keck- stein, “Glenda: gynecologic laparoscopy endometriosis dataset,” in International Conference on Multimedia Modeling , pp. 439–450, Springer, 2019
2019
-
[77]
The saras endoscopic surgeon action detection (esad) dataset: Challenges and methods,
V . S. Bawa, G. Singh, F. KapingA, I. Skarga-Bandurova, E. Oleari, A. Leporini, C. Landolfo, P. Zhao, X. Xiang, G. Luo, et al. , “The saras endoscopic surgeon action detection (esad) dataset: Challenges and methods,” arXiv preprint arXiv:2104.03178 , 2021
2021 arXiv
-
[78]
A. P. Twinanda, Vision-based approaches for surgical activity recogni- tion using laparoscopic and RBGD videos . PhD thesis, University of Strasbourg, Strasbourg, France, Jan. 2017. 11
2017
-
[79]
Multi-task recurrent convolutional network with correlation loss for surgical video analysis,
Y . Jin, H. Li, Q. Dou, H. Chen, J. Qin, C.-W. Fu, and P.-A. Heng, “Multi-task recurrent convolutional network with correlation loss for surgical video analysis,” Med. Image Anal. , vol. 59, p. 101572, Jan. 2020
2020
-
[80]
Single- and multi-task architectures for surgical workflow challenge at M2CAI 2016,
A. P. Twinanda, D. Mutter, J. Marescaux, M. de Mathelin, and N. Padoy, “Single- and multi-task architectures for surgical workflow challenge at M2CAI 2016,” arXiv [cs.CV], Oct. 2016
2016
-
[81]
SV-RCNet: Workflow recognition from surgical videos using recurrent convolutional network,
Y . Jin, Q. Dou, H. Chen, L. Yu, J. Qin, C.-W. Fu, and P.-A. Heng, “SV-RCNet: Workflow recognition from surgical videos using recurrent convolutional network,” IEEE Trans. Med. Imaging, vol. 37, pp. 1114– 1126, May 2018
2018
-
[82]
TeCNO: Surgical phase recognition with multi- stage temporal convolutional networks,
T. Czempiel, M. Paschali, M. Keicher, W. Simson, H. Feussner, S. T. Kim, and N. Navab, “TeCNO: Surgical phase recognition with multi- stage temporal convolutional networks,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2020 , pp. 343–352, Springer Int...
2020
-
[83]
Trans-SVNet: Accurate phase recognition from surgical videos via hybrid embedding aggregation transformer,
X. Gao, Y . Jin, Y . Long, Q. Dou, and P.-A. Heng, “Trans-SVNet: Accurate phase recognition from surgical videos via hybrid embedding aggregation transformer,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2021 , pp. 593–603, Springer Interna- tional P...
2021
-
[84]
LoViT: Long video transformer for surgical phase recognition,
Y . Liu, M. Boels, L. C. Garcia-Peraza-Herrera, T. Vercauteren, P. Das- gupta, A. Granados, and S. Ourselin, “LoViT: Long video transformer for surgical phase recognition,” Med. Image Anal. , vol. 99, p. 103366, Oct. 2024
2024
-
[85]
Surgformer: Surgical transformer with hierarchical temporal attention for surgical phase recognition,
S. Yang, L. Luo, Q. Wang, and H. Chen, “Surgformer: Surgical transformer with hierarchical temporal attention for surgical phase recognition,” in International Conference on Medical Image Computing and Computer-Assisted Intervention , pp. 606–616, Springer, 2024
2024
-
[86]
CholecTriplet2021: A benchmark challenge for surgical action triplet recognition,
C. I. Nwoye, D. Alapatt, T. Yu, A. Vardazaryan, F. Xia, Z. Zhao, T. Xia, F. Jia, Y . Yang, H. Wang, D. Yu, G. Zheng, X. Duan, N. Getty, R. Sanchez-Matilla, M. Robu, L. Zhang, H. Chen, J. Wang, L. Wang, B. Zhang, B. Gerats, S. Raviteja, R. Sathish, R. Tao, S. Kondo, W. Pang, H....
2023
-
[87]
Recognition of instrument-tissue interactions in endo- scopic videos via action triplets,
C. I. Nwoye, C. Gonzalez, T. Yu, P. Mascagni, D. Mutter, J. Marescaux, and N. Padoy, “Recognition of instrument-tissue interactions in endo- scopic videos via action triplets,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2020, pp. 364–374, Springer I...
2020
-
[88]
Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos,
C. I. Nwoye, T. Yu, C. Gonzalez, B. Seeliger, P. Mascagni, D. Mutter, J. Marescaux, and N. Padoy, “Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos,” Med. Image Anal., vol. 78, p. 102433, May 2022
2022
-
[89]
Concept graph neural networks for surgical video understanding,
Y . Ban, J. A. Eckhoff, T. M. Ward, D. A. Hashimoto, O. R. Meireles, D. Rus, and G. Rosman, “Concept graph neural networks for surgical video understanding,” IEEE Trans. Med. Imaging , vol. PP, July 2023
2023
-
[90]
Rendezvous in time: an attention-based temporal fusion approach for surgical triplet recognition,
S. Sharma, C. I. Nwoye, D. Mutter, and N. Padoy, “Rendezvous in time: an attention-based temporal fusion approach for surgical triplet recognition,” Int. J. Comput. Assist. Radiol. Surg. , vol. 18, pp. 1053– 1059, June 2023
2023
-
[91]
Parameter-efficient framework for surgical action triplet recognition,
Y . Li, B. Bai, and F. Jia, “Parameter-efficient framework for surgical action triplet recognition,” International Journal of Computer Assisted Radiology and Surgery , vol. 19, no. 7, pp. 1291–1299, 2024
2024
-
[92]
AdaptiveSAM: Towards efficient tuning of SAM for surgical scene segmentation,
J. N. Paranjape, N. G. Nair, S. Sikder, S. S. Vedula, and V . M. Patel, “AdaptiveSAM: Towards efficient tuning of SAM for surgical scene segmentation,” in Medical Image Understanding and Analysis , Lecture notes in computer science, pp. 187–201, Cham: Springer Nature Switzerland, 2024
2024
-
[93]
Medical transformer: Gated axial-attention for medical image segmentation,
J. M. J. Valanarasu, P. Oza, I. Hacihaliloglu, and V . M. Patel, “Medical transformer: Gated axial-attention for medical image segmentation,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2021, pp. 36–46, Springer International Publishing, 2021
2021
-
[94]
TransUNet: Transformers make strong encoders for medical image segmentation,
J. Chen, Y . Lu, Q. Yu, X. Luo, E. Adeli, Y . Wang, L. Lu, A. L. Yuille, and Y . Zhou, “TransUNet: Transformers make strong encoders for medical image segmentation,” arXiv [cs.CV], Feb. 2021
2021
-
[95]
Encoder-decoder with atrous separable convolution for semantic im- age segmentation,
L.-C. Chen, Y . Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic im- age segmentation,” arXiv [cs.CV], Feb. 2018
2018
-
[96]
UNETR: Transformers for 3D medical image segmentation,
A. Hatamizadeh, Y . Tang, V . Nath, D. Yang, A. Myronenko, B. Land- man, H. R. Roth, and D. Xu, “UNETR: Transformers for 3D medical image segmentation,” in 2022 IEEE/CVF Winter Conference on Appli- cations of Computer Vision (WACV) , pp. 574–584, IEEE, Jan. 2022
2022
-
[97]
nnU-net: a self-configuring method for deep learning-based biomedical image segmentation,
F. Isensee, P. F. Jaeger, S. A. Kohl, J. Petersen, and K. H. Maier- Hein, “nnU-net: a self-configuring method for deep learning-based biomedical image segmentation,” Nature methods , vol. 18, no. 2, pp. 203–211, 2021
2021
-
[98]
UNet++: A nested U-net architecture for medical image segmentation,
Z. Zhou, M. M. R. Siddiquee, N. Tajbakhsh, and J. Liang, “UNet++: A nested U-net architecture for medical image segmentation,” Deep Learn. Med. Image Anal. Multimodal Learn. Clin. Decis. Support , vol. 11045, pp. 3–11, Sept. 2018
2018
-
[99]
U-net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Lecture Notes in Computer Science , Lecture notes in computer science, pp. 234–241, Cham: Springer International Publishing, 2015
2015
-
[100]
Customized segment anything model for medical image segmentation,
K. Zhang and D. Liu, “Customized segment anything model for medical image segmentation,” arXiv [cs.CV], Apr. 2023
2023
-
[101]
LoRA: Low-rank adaptation of large language models,
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” arXiv [cs.CL], June 2021
2021
-
[102]
S-SAM: SVD-based fine-tuning of segment anything model for medical image segmentation,
J. N. Paranjape, S. Sikder, S. S. Vedula, and V . M. Patel, “S-SAM: SVD-based fine-tuning of segment anything model for medical image segmentation,” arXiv [cs.CV], Aug. 2024
2024
-
[103]
Rethinking RGB-D fusion for semantic segmentation in surgical datasets,
M. A. Jamal and O. Mohareri, “Rethinking RGB-D fusion for semantic segmentation in surgical datasets,” arXiv [cs.CV], July 2024
2024
-
[104]
DDA: Dimensionality driven augmentation search for contrastive learning in laparoscopic surgery,
Y . Zhou, H. Badgery, M. Read, J. Bailey, and C. E. Davey, “DDA: Dimensionality driven augmentation search for contrastive learning in laparoscopic surgery,” arXiv [cs.CV], June 2024
2024
-
[105]
Revisiting surgical instrument segmentation without human intervention: A graph partitioning view,
M. Sheng, J. Fan, D. Liu, R. Kikinis, and W. Cai, “Revisiting surgical instrument segmentation without human intervention: A graph partitioning view,” arXiv [cs.CV], Aug. 2024
2024
-
[106]
Contrastive transformer-based multiple instance learning for weakly supervised polyp frame detection,
Y . Tian, G. Pang, F. Liu, Y . Liu, C. Wang, Y . Chen, J. Verjans, and G. Carneiro, “Contrastive transformer-based multiple instance learning for weakly supervised polyp frame detection,” in International Conference on Medical Image Computing and Computer-Assisted Intervention...
2022
-
[107]
Probabilistic representations for video contrastive learning,
J. Park, J. Lee, I.-J. Kim, and K. Sohn, “Probabilistic representations for video contrastive learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pp. 14711– 14721, 2022
2022
-
[108]
Static and dynamic concepts for self-supervised video representation learning,
R. Qian, S. Ding, X. Liu, and D. Lin, “Static and dynamic concepts for self-supervised video representation learning,” in European conference on computer vision , pp. 145–164, Springer, 2022
2022
-
[109]
Parameter- efficient image-to-video transfer learning for action recognition,
J. Pan, Z. Lin, X. Zhu, J. Shao, and H. L. ST-Adapter, “Parameter- efficient image-to-video transfer learning for action recognition,” Preprint at https://arxiv. org/abs/2206.13559 , 2022
2022 arXiv
-
[110]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv [cs.CV], Oct. 2020
2020
-
[111]
TEsoNet: knowledge transfer in surgical phase recognition from laparoscopic sleeve gastrectomy to the laparoscopic part of Ivor–Lewis esophagectomy,
J. A. Eckhoff, Y . Ban, G. Rosman, D. T. M ¨uller, D. A. Hashimoto, E. Witkowski, B. Babic, D. Rus, C. Bruns, H. F. Fuchs, and O. Meire- les, “TEsoNet: knowledge transfer in surgical phase recognition from laparoscopic sleeve gastrectomy to the laparoscopic part of Ivor–Lewis ...
2023
-
[112]
C. R. Johnson and R. A. Horn, Matrix analysis. Cambridge university press Cambridge, 1985
1985
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.