Scientists plan to create a self-learning neural network capable of adapting to different environments
Researchers from MIPT and international research centres have developed Un-ViTAStereo, a new stereo-vision technology that enables robots and autonomous vehicles to perceive the world in three dimensions without blind spots. MIPT’s press service said the algorithm calculates distances to objects without costly LiDAR systems or manual labelling, making it more affordable and versatile.
How Un-ViTAStereo estimates depth
Un-ViTAStereo is trained using the Depth Anything V2 model, which assesses the relative depth of objects from a single image by recognising shadows, perspective and occlusion. This enables the algorithm to retain only predictions that match the “mentor” model’s guidance, improving its accuracy.
The system operates in three stages:
- checking whether every pixel corresponds with the guidance;
- identifying green neighbours for red points;
- creating contours through a disparity-smoothing function.
As a result, the proportion of major errors in the KITTI 2015 autonomous-driving test was reduced to 5%, representing 23% fewer dangerous errors in estimating distances to objects.
Plans for a self-learning neural network
MIPT notes that the current version of Un-ViTAStereo is only a starting point. The scientists intend to develop a self-learning neural network that can adjust to different environments and use precise LiDAR measurements to improve accuracy. The new technology offers broad potential to enhance the safety and functionality of autonomous systems. The research has been published in IEEE Transactions on Circuits and Systems for Video Technology.
Comments
No comments yet. Be the first to comment!
Leave a Comment