Figure announced on August 25, 2026, that its mobile data collection app, Index, is now available on Google Play and the App Store following four months in stealth. The platform pays human contributors to record everyday domestic chores and commercial workplace labor, creating a dedicated data pipeline intended to train the company's Helix AI models.
Contributor Scale and Ingestion Workflow
Figure turned to a direct-to-consumer collection model after concluding that third-party data vendors could not meet its throughput, diversity, or quality standards. Operating in stealth over the past four months, the company reported surpassing 264,000 downloads across 108 countries, supported by more than 44,000 weekly active users. As of the announcement, Figure said it had paid $15 million to contributors, referred to as Creators.
Figure reported that the app was receiving 30 minutes of video every second. Beyond individuals filming their own activities—such as cooking, folding laundry, changing motor oil, or stocking shelves—the platform lets users and businesses book Creators to perform tasks on-site.
To convert raw consumer uploads into structured training data, Figure processes footage through a five-stage ingestion pipeline:
- Filtering: Automated screening evaluates technical, visual, and semantic quality.
- Fraud review: Human analysts perform user-level audits to detect attempts to bypass automated filters.
- Deduplication: Video segments are embedded to compute similarity against existing footage, discarding clips exceeding a similarity threshold.
- Rebalancing: Task quotas and embedding-based clusters rebalance retained footage based on how closely submissions align with targeted tasks and diverse variations.
- Annotation: The pipeline generates hierarchical text captions linked to each accepted episode.
Figure claims that for every 1,000 hours collected, Index captures 373 unique tasks, 1,146 unique manipulated objects, and 116 unique environments.
Human Video Diversity Versus Robot Actuation
Figure frames generalization in robotics as an empirical data problem, arguing that internet text and synthetic data lack the physical nuances of real-world object interaction. By crowdsourcing recordings across international households and workplaces, Figure aims to capture long-tail visual and procedural variation that is difficult to script.
However, passive video captured by human participants differs substantially from the sensorimotor control policies, embodied dynamics, and actuation constraints required to operate physical humanoids. Figure claimed that internal results validate its thesis that crowdsourced human data drives generalization in Helix, adding that it plans to release detailed findings soon. The announcement did not provide robot task-success benchmarks demonstrating the effect of Index data.
