Before a machine can act, it has to agree with us about what it is looking at. Cameras and LiDAR return points and pixels — not "a child", "a kerb", or "a doorway". Somebody has to draw that line first.
Perception is where the physical world is turned into something a policy can reason over. A LiDAR sweep is a few hundred thousand coordinates with intensity values; a camera frame is a grid of brightness. Neither contains the word "pedestrian". That meaning is supplied by people, once, carefully — and every model trained downstream inherits both the accuracy and the mistakes of that first act of interpretation.
What this needs from people
The research
Talks & demonstrations
Where HSV fits
This is work HSV already delivers. Our 3D LiDAR annotation line labels continuous-frame point clouds with 3D boxes and camera context, including the missing and distorted returns most pipelines quietly drop. Our parking-slot work is the same discipline at tighter tolerance: bounding boxes plus inner corner points, with attributes chosen by slot and surface type.
See the delivery line