AI adoption is no longer the only signal leaders should watch. The Stanford AI Index reports that 88% of organizations use AI in at least one business function, while McKinsey's 2025 survey found that only about one-third of organizations had begun scaling their AI programs.[1] [2] The studies use different samples and methods, so the figures should not be read as a direct comparison. Together, however, they point to an important market question: what separates access to AI from dependable production use?
Adoption and scale are different market signals
Adoption shows that AI has entered the enterprise agenda. Scale shows whether organizations can redesign workflows, connect systems, maintain data and evaluate performance well enough to use AI repeatedly. McKinsey reports that high-performing organizations are nearly three times more likely than others to redesign workflows, and it identifies data infrastructure and human validation practices among the factors associated with value.[2]
This distinction supports a measured position. Investment in models remains necessary and is led by organizations with significant resources and technical capability. The operational opportunity sits around those models: preparing the data they need, testing how they behave and maintaining evidence as the environment changes.
Physical AI makes the production gap visible
A laboratory can hold a camera angle steady, control lighting and define expected objects. A live store, warehouse or outdoor environment cannot. Products move. Layouts change. Sensors drift. Equipment wears. Weather intervenes. People improvise.
Gartner's research on physical-AI scaling identifies data scarcity, sim-to-real gaps and embodiment gaps as barriers that delay development and increase cost.[3] Better models can reduce some errors, but they still need evidence from the environments where they will operate. A missing condition cannot be evaluated, and an unclear definition cannot produce dependable ground truth.
A production example: Amazon Vulcan
Amazon describes Vulcan as its first robot with a sense of touch. The company reports that it can handle approximately 75% of the item types stored in its fulfillment centers at speeds comparable to employees. Amazon also says the system was trained using physical information such as touch and force feedback, together with thousands of real-world examples.[4]
The most useful detail is that Vulcan is designed to know when it cannot handle an item and ask a person for help. That is not evidence that one approach applies to every warehouse. It is a company-reported example of a broader production principle: the operating system needs to detect uncertainty, route it appropriately and learn from the result.
What the data operation must provide
The data operation connects capture, ground truth, specialist review, evaluation and model feedback. It decides which examples represent the intended environment, how those examples are annotated, how disagreement is handled and which failures carry the greatest operational cost.
This is where Trinovation's position is most credible. The company does not need to argue that model investment is misplaced. It can show that model performance depends on disciplined annotated-data operations and evaluation evidence that survive contact with the real world.
Give the layer clear ownership
The data layer often spans product, engineering, model development, quality and external operations. That distribution is understandable, but it creates risk when no one owns the end-to-end learning loop. One team can optimize collection while another uses a different definition of quality, and the gap may not appear until deployment.
Leaders should define accountable owners for taxonomy, evaluation, data-version decisions and operational feedback. Ownership does not require one department to perform every task. It requires clarity about who resolves conflicts and who decides when evidence is strong enough to change the product.
A practical governance rhythm can include a recurring review of new edge-case patterns, gold-set changes, environment shifts and unresolved escalations. The review should connect operational evidence to a specific decision rather than become a general status meeting.
Design for expansion into new environments
Every new store format, facility, device, geography or customer introduces the possibility that previous assumptions will fail. Expansion planning should therefore include a data-readiness assessment: what will change, which examples are missing, how evaluation will be localized and what specialist knowledge is required.
This is where a dependable data layer becomes a commercial advantage. The organization can enter a new environment with a repeatable adaptation process instead of rebuilding definitions and quality controls after problems appear. Faster adaptation comes from operational learning, not from assuming the model will generalize automatically.
A readiness test for the data layer
A team is ready to scale when it can explain how real-world changes enter the data process, how uncertain cases are resolved and how the resulting evidence changes training or evaluation. It should also be able to reproduce an important decision by tracing the data, definition, reviewer action and model or workflow response.
If those connections depend on individual memory, the layer is not yet dependable. The next investment may be less about collecting more examples and more about strengthening ownership, versioning, evaluation and feedback. That operating foundation allows future data to create value faster.
Sources & Review Note
- Stanford Institute for Human-Centered AI, AI Index Report — enterprise AI adoption.
- McKinsey & Company, The State of AI (2025 survey) — scaling, workflow redesign, and value factors.
- Gartner — research on physical-AI scaling: data scarcity, sim-to-real gaps, embodiment gaps.
- Amazon — company announcement on the Vulcan warehouse robot.
Stanford and McKinsey use different samples and should not be treated as one continuous dataset. Amazon figures are company-reported. No internal Trinovation observation is presented as evidence.
Building AI that has to work in the real world?
Trinovation runs the annotated-data operations behind production computer vision and spatial AI — ground truth, IAA audits, and evaluation evidence included. First delivery in 1–2 weeks.
Book a 20-Min Call → Start a Free 500-Frame Pilot