Everyone keeps asking which robot foundation model will win. The less glamorous answer may be whichever one has the best choreographed data exhaust. Humanoid robots do not learn to fold towels, dodge pallet jacks, or grasp a slippery wrench by inhaling the web and hoping physics feels generous. They need recorded reality, which is annoying because reality has terrible API documentation. China’s National Data Administration is now stepping into that mess. The agency said it will develop standards for embodied AI training data and guide local authorities, according to Bloomberg, landing just ten days after seven companies asked for public data infrastructure and common standards, according to The Next Web. That timing is the story. The robots are not just waiting for a bigger model brain, they are waiting for a national data pipeline with labels, governance, and enough consistency that training does not become artisanal chaos in a lab coat. ## The ask became policy machinery fast, according to The Next Web The Next Web reported that the National Data Administration’s move followed a 10 September meeting chaired by Liu Liehong, the agency’s head, with research institutes, technology companies, and humanoid robotics organizations attending. Ten days earlier, the same agency met seven companies that requested public data infrastructure for embodied AI and common data standards, according to the same report. That is unusually direct feedback loop energy, at least by policy standards, where urgency often moves like a Roomba trapped under a sofa. Bloomberg reported that the National Data Administration plans to strengthen planning and development guidance for the sector while supporting companies increasing investment in data resources. TradingView News, summarizing the initiative, said the priorities include formal standards for how embodied AI data is collected, labeled, stored, and shared, plus guidance for local governments on data governance. In less bureaucratic English: the country is trying to standardize the messy middle between a robot touching the world and a model learning from that contact. ## The bottleneck is hours of reality, The Next Web reports The Next Web cites the China Academy of Information and Communications Technology’s estimate that embodied AI foundation models need roughly 10 million hours of real training data. The same report says only 100,000 to 1 million hours of high quality data exist worldwide. That gap is not a rounding error, it is the Grand Canyon wearing a motion capture suit. This is why embodied AI is different from text models in a very expensive way. Text is abundant, imperfect, and legally spicy, but it exists at planetary scale. Robot interaction data is harder because it must capture physics, spatial awareness, sensor readings, action outcomes, and human movement in settings where gravity keeps filing bug reports. China is already building supply, according to The Next Web, which reports that more than 70 embodied AI training grounds are operating, mostly for industrial manufacturing, with 46 more planned. If those facilities produce data under shared standards, private model developers get more than raw hours. They get comparable hours, which is the difference between a dataset and a junk drawer with timestamps. ## Standards are becoming the stack, TradingView News says TradingView News describes embodied AI as systems that perceive and physically interact with their environment, from warehouse robots to surgical arms to humanoids. That breadth matters because a data standard has to survive wildly different machines, sensors, tasks, and safety expectations. A robot arm in a factory and a humanoid in a demo booth both produce machine generated data, but pretending those streams are magically interchangeable is how you get benchmark theater with better lighting. The AI Insider reported that China issued its first national standard system for humanoid robots and embodied AI, unveiled on February 28 in Beijing. That framework was organized into six components, including basic commonality, brain like and intelligent computing, limbs and components, complete machines and systems, application, and safety and ethics, and was drafted by more than 120 institutions under the Ministry of Industry and Information Technology’s technical committee. The same report said more than 140 domestic manufacturers released over 330 humanoid models in 2025, which helps explain why standards are not paperwork garnish here. They are traffic lights before everyone sprints into the intersection carrying servo motors. ## Europe has access rules, not production pipes, says The Next Web The Next Web contrasts China’s buildout with Europe’s Data Act, which governs access to machine generated data but does not create the infrastructure to produce it. The report says most actions from the EU’s Data Union Strategy, including planned data labs, had not been carried out ten months after publication. Access rules are useful, but they are not data collection rigs, calibrated sensors, shared schemas, or training grounds where robots repeatedly fail at real tasks until they become slightly less embarrassing. The demand for national infrastructure is not new inside China either. Yicai Global reported that Haier Group chairman and chief executive Zhou Yunjie called for national level open innovation platforms for embodied intelligence, a dedicated national data program, and an industrial grade standards framework spanning design, manufacture, testing, and deployment. That earlier ask rhymes with the seven company request The Next Web reported, suggesting industry sees data infrastructure as a shared dependency rather than a nice policy accessory. For AI builders, the lesson is wonderfully unsexy and therefore probably important: watch the data layer. Model releases will still get the glossy demos, the applause, and the vaguely humanoid jazz hands. But if physical AI scales, it may be because somebody standardized how robots record the world, label their mistakes, and share enough reality for models to stop learning from vibes. The robot revolution will not arrive with a monologue, it will arrive as a schema migration. ## Sources - China’s data regulator plans standards for embodied AI, ten days after industry asked

Sources