Realset AI, a real-world data lab based in San Jose, and Flatkey, an AI infrastructure platform, have raised $10 million in Series A funding. The investment will expand Realset's capture network of real workplaces and studio environments while growing its pool of expert demonstrators and domain experts. It will also support open benchmarks that measure whether AI policies work outside the lab.
Why Real-World Data Is Needed
Frontier labs and robotics companies have largely consumed the text available on the internet, making physical human activity the next major data frontier. Tasks like folding laundry, loading a dishwasher, packing an order, and handling a customer return do not exist on the web and cannot be reproduced through simulation alone. Realset argues that every dataset must begin with a real person performing a real task in a real place.
Funding Strategy and Expansion
The Series A round will support the expansion of Realset's physical capture network and its expert demonstrator community. Funding will also help the company build reinforcement learning environments that mirror real commercial operations such as e-commerce, customer support, and logistics. Flatkey, which provides developers access to more than 100 official AI models and more than 1,000 AI tools through one key, participated in the round.
Three Ways to Produce Ground Truth
The company's Realset Body service captures embodied data for physical policy models using egocentric and third-person video of skilled workers in homes, kitchens, warehouses, and light assembly lines. Realset Field builds LLM and agent training data from reinforcement learning environments modeled on real e-commerce, customer support, logistics, and manufacturing workflows. Realset Judge provides expert evaluation, failure diagnosis, and continuous monitoring for agents in production, with targeted fix data for common failure modes.
Workspace and Quality Assurance
All three services operate through the Realset Workspace, where domain experts fluent in six languages complete real tasks, attach evidence and screen recordings, and pass independent quality review before records are approved. Approved records are exported as JSONL with provenance that customers can verify. This structure combines task execution with documented evidence to maintain consistent data quality.
Open Benchmarks on Real Tasks
Realset publishes open benchmarks to measure whether AI policies succeed outside the lab. The Realset Household Manipulation Bench will evaluate open-source vision-language-action policies such as OpenVLA and GR00T on folding, loading, sorting, and wiping tasks captured in real kitchens and laundry rooms. Results are expected in the fourth quarter of 2026, with additional benchmarks planned for light assembly and commerce operations.
Leadership Perspective
Hunter Guo, founder of Realset AI, said the easy data is gone and that the physical world cannot be scraped. He explained that the work requires placing a camera on a skilled person doing real tasks, structuring what they did, and checking it with people who know the job. Guo described the challenge as a capture and quality problem rather than a labeling problem.
By combining expert human demonstrations, real-world reinforcement learning environments, and domain expert evaluation, Realset AI aims to supply the ground truth data that frontier models and embodied agents need. The planned open benchmarks will give the broader field a way to compare policies on real household and commercial tasks. With fresh Series A capital, the company is moving to expand its capture network and build publicly available evidence of how AI performs outside controlled settings.