InstEditSeg: Instruction-Driven Image Editing for Polyp and Skin Lesion Segmentation
Ziquan Liu, Zhewei Zhu, Xuyang Shi
Abstract
Accurate segmentation of polyps and skin lesions is pivotal for clinical diagnosis, yet existing methods struggle with low contrast, ambiguous boundaries, and cross-domain distribution discrepancies. Discriminative networks and most diffusion-based segmentation approaches predict standalone binary masks, leaving the visual priors of large-scale pretrained generative models largely unexploited. We propose InstEditSeg, a unified generative framework that reformulates medical segmentation as an instruction-driven image editing problem. Instead of emitting a mask, the model renders a color-coded overlay on the original image, conditioned on a textual instruction, so that the edited output aligns with the natural image distribution learned by latent diffusion models and mitigates the domain gap between natural and medical imagery. To recover fine anatomical structures, we introduce DINOv3 as an auxiliary visual encoder and a DINO Feature Guidance Block that builds a multi-scale feature pyramid. The pyramid is fused into the diffusion U-Net by channel concatenation and zero-initialized convolution so that hierarchical discriminative priors can be injected without perturbing the pretrained weights. A dual-branch classifier-free guidance strategy requiring only two forward passes per denoising step reduces inference cost. On polyp and skin lesion benchmarks the framework achieves accuracy competitive with strong discriminative baselines, and it further demonstrates concrete advantages of the generative formulation: notably better cross-domain generalization on unseen data, more complete multi-lesion segmentation, instruction-conditioned task control, and sampling flexibility. We also analyze the strengths and limitations of the paradigm, including its color sensitivity and unsupported attribute-conditioned selection. Code is available at: https://github.com/wincharm001/InstEditSeg.
Create a lesson
Related papers
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
Junchao Huang, Guian Fang, Shengju Qian et al.
Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image Generation
Yutong Liu, Nan Huang, Xu Cao et al.
PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud Segmentation
Yu Tian, Xintong Jiang, Jan Franklin Adamowski et al.
MuyBridge: Mobile Human Center-of-Mass Estimation from Monocular Video via Sparse Fusion
Aidan Bradshaw, Marco Giordano, David Rode et al.
RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation
Xiaolei Lang, Ze Kang, Zehao Huang et al.
Efficient All-in-One Weather Restoration using Spectral Harmonization
Paula Garrido-Mellado, Daniel Feijoo, Yuning Cui et al.