What if your robot could understand any object you describe, just...

RADIO-ViPE builds a 3D map from raw monocular video that you can query with natural language.
(1/4)
A foundation model (RADIO) extracts dense "meaning vectors" per pixel, then they reuse those same vectors three ways: improving optical flow on blank surfaces, adding a semantic loss into the geometry optimizer, and connecting similar keyframes in the factor graph.
