FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
Maya Moriya, Sigal Raab, Yael Vinker, Tali Dekel
Abstract
We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions. To translate visual content into folding programs, we define a parametric space that consists of the paper's geometry and a set of parametric folding actions. Unlike models that predict static crease patterns, our agent operates sequentially and possesses the ability to re-plan its actions, effectively mitigating the compounding errors inherent in multi-step folding. Our approach takes a step toward closing the gap between human origami knowledge, which is primarily shared through unstructured visual demonstrations, and computational methods, which typically rely on structured, parametric representations such as a crease pattern or an executable parametric plan. We evaluate our approach on PurelandFold, a newly curated benchmark of diverse Pureland origami videos with ground-truth geometry and action labels. Our results demonstrate that by combining VLM reasoning with a set of specialized tools and physical simulation, we can successfully transform unstructured visual demonstrations into executable, physically plausible folding procedures.
Create a lesson
Related papers
Instance-Guided Report Anchoring for Text-Free 3D Abnormality Segmentation in Chest CT
Zhenyu Bu, Haoyan Ding, Chushu Shen et al.
Unmasking Face Embeddings: Reading, Rendering and Naming with Foundation Models
Fizza Rubab, Yiying Tong, Arun Ross
SlideMix: Enhancing Whole Slide Image Analysis via Multimodal Shuffling
Chad Wong, Sicheng Chen, Tianyi Zhang et al.
Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Vida Adeli, Soroush Mehraban, Jacob Rommann et al.
Where Should Experience Live? Hierarchical Hebbian Memory for Continual Vision Transformers
Mohammed Yusuf Mujawar, Noorbakhsh Amiri Golilarz
TRUST: Threshold-Recalibrated Uncertainty-Safe Training for Certified Dismissal in Breast Cancer Screening
Parham Hajishafiezahramini, Matthew Hamilton, Edward Kendall et al.