Embodied AI
The head01 / 06

Perception

Before a machine can act, it has to agree with us about what it is looking at. Cameras and LiDAR return points and pixels — not "a child", "a kerb", or "a doorway". Somebody has to draw that line first.

Perception is where the physical world is turned into something a policy can reason over. A LiDAR sweep is a few hundred thousand coordinates with intensity values; a camera frame is a grid of brightness. Neither contains the word "pedestrian". That meaning is supplied by people, once, carefully — and every model trained downstream inherits both the accuracy and the mistakes of that first act of interpretation.

What this needs from people

  • Three-dimensional boxes drawn across continuous frames, so an object keeps its identity as it moves.
  • Camera context matched to the point cloud, so the two views describe one scene.
  • Edge cases handled deliberately — missing returns, distorted points, partial occlusion.
  • Precise geometry where a few centimetres decide whether a manoeuvre is safe.

Talks & demonstrations

The people building it, in their own words.

CVPR Workshop on Autonomous Driving

Keynote — Andrej Karpathy, Tesla

PyTorch

PyTorch at Tesla — Andrej Karpathy

Where HSV fits

This is work HSV already delivers. Our 3D LiDAR annotation line labels continuous-frame point clouds with 3D boxes and camera context, including the missing and distorted returns most pipelines quietly drop. Our parking-slot work is the same discipline at tighter tolerance: bounding boxes plus inner corner points, with attributes chosen by slot and surface type.

See the delivery line
Himalayan Silicon Valley
Himalayan Silicon Valley Pte. Ltd.
Lumbini · Kathmandu · Global
The AI and technology arm of The Promised Group. Building compute, products and a trained workforce out of Nepal, for the markets that buy them.
© 2026 Himalayan Silicon Valley · Nepal. All rights reserved.