The policy is the part that decides. Capability alone does not make that decision good — a model can be fluent, accurate and still choose the wrong action. Preference and safety judgement are human signals, and there is no synthetic substitute for them yet.
The technique that made modern models useful was not a bigger network — it was asking people which of two outputs was better, thousands of times, and training on the answer. That method now reaches into robotics: the same preference signal that shapes how a model writes is being used to shape how a machine moves. It means the quality ceiling of an embodied system is set by the consistency of the people judging it.
What this needs from people
The research
Talks & demonstrations
Where HSV fits
Image preference evaluation is one of our delivery lines, and the review system behind it is the point: every project runs against a written QA handbook with a double-review pass, so a judgement made in week one still holds in month six. That is the same machinery a preference-trained policy depends on.
See the delivery line