Video alone does not always provide the structure required by downstream robotics models.
IndiVillage can add an annotation and metadata layer that makes manipulation episodes easier to search, train on, compare, and evaluate.
A multi-step task can be segmented into meaningful actions such as reach, grasp, lift, rotate, place, insert, open, close, hand over, or release. Object states can be linked to those actions, and important moments such as failed attempts, corrections, or task completion can be marked consistently.
Depending on the model requirement, workflows can also incorporate hand and object localization, keypoints, temporal boundaries, interaction labels, task-stage labels, and outcome metadata.
This creates a clearer connection between what happened in the demonstration and the learning signal the model needs.