AGRO-SUVIDE:
AGentic RObotics for SUrgical VIscoelastic DEbridement

Anonymous Authors

Abstract

Augmented dexterity has the potential to reduce the fatigue experienced by surgeons during repetitive surgical tasks. In this paper, we propose the first AGentic RObotics framework for SUrgical VIscoelastic DEbridement (AGRO-SUVIDE), the repeated removal of small fragments attached to a viscoelastic substrate. Leveraging the self-improving and coding capability of agents, AGRO-SUVIDE adopts a modular framework. Specifically, the demonstration analysis module automatically identifies recurring skills from a single expert demonstration, using both visual and kinematic information. The construction module then builds each skill, either as a procedural model-based skill the agent codes against a scaffolded library or as a model-free policy-based skill. At runtime, the monitoring module composes the skills into a loop-style graph sized to the number of fragments it observes, then verifies pre- and post-conditions of each skill to decide whether to advance or retry. We evaluate AGRO-SUVIDE through 340 physical trials on the da Vinci Research Kit (dVRK). AGRO-SUVIDE achieves an average single-fragment removal success rate of 85%, completing consecutive three-fragment removal at 60% and at 95% with one human intervention. It further generalizes to unseen five-fragment scenarios with an average success rate of 80% for single-fragment removal.

AGRO-SUVIDE concept: agents identify modular skills from one expert demonstration and a monitor verifies their pre- and post-conditions at runtime
AGRO-SUVIDE: Agents automatically identify and improve modular skills from one expert demonstration, build each as model-based or model-free, and a monitor executes them at runtime. Here the monitor confirms that the model-based skill “lift” has met its post-condition of Δz ≥ 0.029 m, and holds the model-free skill “cut” running until its post-condition is met.

Research Questions

Can the self-improving capability of agents exploit the repetition inherent in debridement
to identify and improve modular skills automatically?

How do we ensure the identified skills are executable and chain reliably?

The AGRO-SUVIDE Framework

AGRO-SUVIDE pipeline: demonstration analysis, construction, and monitoring modules over a shared set of identified skills

AGRO-SUVIDE adopts a modular design. Given one expert demonstration, the demonstration analysis module exploits the repetition in surgical debridement to identify and improve modular skills with explicit pre- and post-conditions defined over vision and kinematics. The construction module then reads those conditions and builds each skill to satisfy them, either as a procedural model-based skill coded against a scaffolded library or as a model-free policy-based skill trained on data segments bounded by the same conditions. At runtime, the monitoring module composes the loop-style graph sized to the number of fragments it observes, and verifies the conditions to decide whether to advance or retry.

Demonstration Analysis Module

The module exploits repetition in surgical debridement to identify and improve modular skills automatically from a single expert demonstration. Its pipeline of pre-processing, semantic merge, and improvement under boundary and repetition constraints returns skills with explicit pre- and post-conditions defined over vision and kinematics.

Skill graph growing from one expert demonstration

Construction Module

Builds each identified skill against its identified conditions. Procedural model-based skills are coded by the agent over a scaffolded library exposing the dVRK control API and foundation models such as SAM3 for segmentation and RAFT-Stereo for depth. Skills that a model-based implementation cannot satisfy reliably are re-implemented as model-free policy-based skills trained with ACT, using a compositional data collection protocol that bounds every recorded segment by the same identified conditions so model-based and model-free skills chain without adjustment.

Monitoring Module

At runtime, pairs the skills into phases and composes a loop-style graph sized to the number of fragments it observes. It verifies the pre- and post-conditions of each skill: execution advances when the post-condition is satisfied; otherwise the monitor samples a new pose satisfying the pre-condition and retries the skill.

Results

We evaluate AGRO-SUVIDE through 340 physical trials on the dVRK. On three-fragment debridement, it reaches an 85% single-fragment success rate, compared to 42% for the model-based baseline and 18% for the model-free ACT baseline. It completes consecutive three-fragment removal at 60%, and at 95% with one human intervention, with the highest throughput of 65 fragments per hour. On unseen five-fragment scenarios, it generalizes with an 80% single-fragment success rate.

Video Demonstrations

Comparison: Three-Fragment Debridement

Model-Based

10× Speed

Model-Free

10× Speed

AGRO-SUVIDE

10× Speed


Failure Modes

Model-Based

Incomplete cut due to pose-dependent kinematic error, misgrasp due to glare of the substrate.

10× Speed

Model-Free

Policy drifts when deviations such as misgrap or miscut happen.

10× Speed

AGRO-SUVIDE

Incomplete cut, a thin residual strand escapes the monitor.

10× Speed


Generalization: Unseen Five-Fragment Scenario

Three-Fragment

10× Speed

Five-Fragment (Unseen)

10× Speed

BibTeX

@inproceedings{anonymous2026agrosuvide,
  title={AGRO-SUVIDE: Agentic Robotics for Surgical Viscoelastic Debridement},
  author={Anonymous Authors},
  booktitle={Under Review},
  year={2026}
}