On the production line or in the cloud?
In the same week that OpenAI released its most powerful model to date, a production line for a humanoid robot – which carries its model on board – began operations in Guangzhou. Both reports are seen as evidence of the same progress. In fact, they describe two different architectures, and for production, the choice between them is the real decision.
14 Sep 2026Share
On 3 September, OpenAI unveiled GPT-6 Astra; on 8 September, XPeng commissioned the production line for its humanoid IRON in Guangzhou. Both events were interpreted as evidence of just how rapidly artificial intelligence is currently advancing. For manufacturing companies, however, the more interesting observation is a different one: the two systems place intelligence in completely different domains, and the consequences of this are operational, not philosophical.
The machine thinks for itself
XPeng manufactures IRON on a production line whose core processes, according to the company, are over 80 per cent automated and designed to automotive quality standards. The robot has 76 degrees of freedom – 21 per hand – and is equipped with three in-house Turing AI chips offering a computing power of up to 2,250 TOPS. XPeng’s Foundation Model for Physical AI runs directly on this hardware within the device. Series production is scheduled for the end of the year, with commercial deliveries in China and abroad set to follow in 2027.
The key figure is not the number of degrees of freedom, but the on-board computing power. XPeng is not building a robot that requires a connection to function. It carries its own model, and that is a deliberate and costly decision: silicon area, cooling and energy within the device cost money and add weight that is then lacking elsewhere.
The data centre thinks for itself
Astra represents the alternative model. OpenAI describes the model as state-of-the-art in computer operation, software engineering, scientific and professional work, and multi-stage workflows, where it adheres more reliably to task boundaries and user intentions than the previous generation. According to the company, it is the result of the largest training run to date, conducted for the first time on more than 100,000 GPUs. It processes over one million tokens of context. And in OpenAI’s internal risk assessment, it reaches the critical threshold in the area of cyber security, which is why the publicly available version refuses to carry out certain security-related tasks.
These capabilities exist exclusively in the data centre. They cannot be installed on a machine, and they are billed per operation: the API price is ten US dollars per million input tokens and 50 US dollars per million output tokens. This is not a one-off purchase, but a variable cost that increases with usage.
Why this is not a question of performance
The instinct to compare both approaches on a scale and ask which is smarter is misleading. They differ in four characteristics that carry more weight in production than any benchmark.
Response time: A gripping operation, collision avoidance or seam tracking cannot tolerate network latency. Anything decided within a machine’s cycle time must be decided locally. Anything decided within the cycle of a shift or a job can wait.
Data outflow: A model in the data centre sees only what is sent to it. When it comes to design status, recipes and test reports, this is a contractual and regulatory issue, not a technical one. A model within the device sees everything and passes nothing on.
Failure behaviour: A locally computing system degrades when the connection is lost; a cloud-based one remains operational. For a production line that cannot afford downtime, this is the crucial difference, and it determines what is actually suitable as a fallback solution.
The cost model: On-device intelligence is an investment with a known cost and a long depreciation period. Cloud intelligence is a running cost that grows in proportion to usage. Anyone wishing to scale an application should know which of the two cost curves they are following.
A useful decision rule
For most businesses, the answer will not be one location, but two. It makes sense to draw a distinction based on cycle time: anything that needs to be decided within milliseconds to seconds belongs on the device, using a model small enough for the available hardware. Anything decided within minutes to hours – such as planning, diagnostics, documentation, quotation calculation, and fault analysis across multiple systems – belongs in the cloud, because that is where the capabilities lie that cannot be afforded locally.
The interface between these two levels is the aspect most frequently underestimated in practice. It determines which data leaves the plant, how often the local model is updated, and what happens if the upper level fails to respond. Those who define this interface clearly can switch providers at both levels. Those who leave it to a single provider have effectively committed to that provider for both levels.
Guangzhou and San Francisco demonstrated in the same week that both approaches are viable. The industry’s task is not to choose one over the other, but to draw the line between them itself before a supplier does.
Related Exhibitors
Interested in news about exhibitors, top offers and trends in the industry?
Browser Notice
Your web browser is outdated. Update your browser for more security, speed and optimal presentation of this page.
Update Browser