Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation
Lili Wang, Jing Li, Xiaowen Sun, Xiangyu Hu, Zhuangzhuang Gu, Jian Liu, Srihari Nelakuditi, Yan Tong
Abstract
Episodic test-time adaptation resets a frozen segmenter to source weights M0 on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, how far to adapt, with an irreducibly per-case one, whether this case should be adapted at all. Cohort means hide that decision: on cross-vendor cardiac MRI the mean ΔDice from adaptation is statistically indistinguishable from zero while 58.7% of cases are individually made worse. We quantify this harm as harmful accepted area (HA), the harmful fraction of the edited area a controller deploys. Held-out tuning gives a stronger baseline than a fixed horizon, but the budget it selects transfers on neither of the two main medical benchmarks, and no global budget can condition on the case. We show that prediction fragmentation---the disagreement geometry between M0 and the adapted mask Mk---predicts HA with no labels or extra backward passes at decision time, comparably on three benchmarks (Spearman ρ 0.50--0.60), at a quarter of gradient-norm's latency. A case-level router built on it cuts HA from 0.228 to 0.139 on a benchmark that took no part in its design, with the design frozen and only cut-points recalibrated there. On the cardiac benchmark the design was selected on, the router cuts HA from 0.129 to 0.013 at matched Dice and 1.10 deployed updates, against the retrospective-best budget found post hoc on evaluation labels, and reduces that 58.7% to 20.0%, an upper bound we quantify. Where the retained cases are not net-helped (as on prostate), the router still cuts HA but concedes accuracy, a boundary we report. Thresholds are fit once on a labeled split disjoint from evaluation; decisions use no labels or gradients. The template ports across architecture and domain (nnU-Net, Cityscapes) with coordinate, thresholds and per-bucket actions instantiated per domain.
Create a lesson
Related papers
FlowSGS: Improving Flow Matching Priors for Inverse Imaging with Stochastic Interpolants
Tianao Li, Xinhui Qian, Emma Alexander
FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents
Dennis Rotondi, Abdelrhman Werby, Kai O. Arras
Earth Surface Immune System for Rapid Monitoring of Unknown Anomalies
Jingtao Li, Qian Zhu, Xinyu Wang et al.
PROVIA: Procedure State Tracking for Online Mistake Detection in Egocentric Videos
Di Wen, Kailun Yang, Jimmy Weissert et al.
Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
Yulong Chen, Ziqian Zhang, Haoyu Zhang et al.
PhGS: Post-Hoc Pruning and Refinement of Single-View Feed-Forward 3D Gaussian Reconstructions
Rinto Yagawa, Han Cheng, Dieter Schmalstieg et al.