The best VR headsets just have cameras on them. That's where the hand tracking comes in. So why can't I just have more cameras around me and then display on a screen? Hacky but can work. Could even be better in the sense I have more gesture control. Hard problem would be syncing the camera views to not clash with each other.
Can I not just use MediaPipe? Why not use 2 cameras? Soso do-able.
I am becoming more MediaPipe pilled, bc the goggles are just for my viewing. And if I post my stuff most people aren't gonna be viewing in 3D anyway. And actually MediaPipe is just the hand tracking. I can render objects via video or live! Or even better CONTROL OBJECTS overlayed a videofeed without my hands, eg a controller.
I want to get an AR headset and start making example interfaces. But for some reason it doesn't feel that high ROI. Like people will vibecode demos for these things in time, is writing / thinking through the shape of things a better use?
See Input devices and The hyperobject.