Dynamic visual–inertial odometry (VIO) requires reliable suppression of motion-corrupted measurements, yet prior
semantic-assisted approaches depend on category-limited segmenters and degrade under partial occlusion. Promptable
foundation segmentation models offer category-agnostic dynamic parsing, but their effectiveness in VIO depends
critically on the temporal stability of input prompts — a factor largely overlooked in existing
pipelines. When prompts derived from raw detection are jittery or intermittent under occlusion, the resulting masks
flicker across frames, destabilizing geometric estimation.
We propose STAG-VIO, which formulates dynamic robustness as a
perception-to-geometry interface stabilization problem. We introduce uncertainty-adaptive
multi-object tracking that models prompt generation as state estimation with bounded noise adaptation, producing
temporally coherent box prompts. These stabilized prompts drive a lightweight foundation segmenter whose masks
undergo geometry-oriented morphological refinement to establish conservative safety margins. A constraint-budget-aware
feature redistribution strategy preserves well-conditioned static measurements when dynamic regions dominate the view.
Experiments on VIODE and OpenLORIS-Scene show consistent gains over state-of-the-art baselines. Ablation confirms
that prompt stabilization is the single most impactful component, reducing trajectory error by up to 83%.