Home About Insights Poems Photos

Insights

Rethinking Unit Economics for the Physical AI Era

AI Can Build the Model. It Can't Broker the Trust Between Capital and Organization. That’s Strategic Finance’s Job.

Robots. When Are They Going to Make My Breakfast?

← Insights

Rethinking Unit Economics for the Physical AI Era

Those in finance, and even those who just touch it, often hear the term "unit economics." The number of times I have had to explain unit economics is staggering. And quite often I used different ways to approach the question. This inconsistency in how I explained unit economics prompted me to sit down and think through what unit economics actually is, why it is used so frequently in finance, and what is different about it in the Physical AI space.

What is unit economics?

Unit economics is a thing or an activity that generates value for its user. A share of that value is captured by a company as revenue. Delivering that value to the user requires the company to deploy resources that either exceed the captured share of value (generating a loss per unit) or, what companies hope for, fall significantly below that share in value (generating profits per unit). It's not the same as the gross or net profit you see in financial statements, as it doesn't account for the noise of overheads, one-time costs, R&D investment, sales force, etc. It's a basic foundational atom of economic value creation at the company that is tied to the nature of the company's product. Different products within the same company will have different unit economics. It might be a commercial transaction processed, a robotaxi ride, a customer acquired, a traffic camera, a nuclear microreactor, an oil well, an AI chatbot query, a subscription seat, a mile of freight delivery, a shipping container; the list goes on. Additionally, the same activity can be looked at from different angles, resulting in different unit economics. For instance, unit economics for commercial freight delivery can be a mile, a truck, or a trip. Regardless of the unit economics a company chooses, consistent usage, appropriate revenue and cost allocation, and their inclusion in actual decision-making are crucial.

Unit economics helps companies analyze multiple questions:

1. Do we deploy resources efficiently?

2. Are we able to capture maximum value from our product?

3. Does scaling hurt or help our profits?

A business with healthy unit economics gets stronger as it grows. A business with broken unit economics loses money faster if there is no realistic plan in place to improve it. The aggregate P&L will eventually tell you which one you have. The beauty of unit economics is that it tells that story ahead of time.

Unit economics of a physical product.

In the classic world of physical goods, the concept of unit economics is quite simple. It's commonly the physical good itself (a thing that exists in the physical world as a physical object). For example, a car, a microwave, or a goat. The value is derived by taking its consumer price and deducting the marginal cost to make it and deliver it into the hands of the end consumer, including raw materials, direct labor, freight, and other variable and directly attributable costs. The unit is a discrete object created, delivered, and consumed at a point in time. Costs are primarily COGS that scale but likely improve on a per-unit basis with volume. Companies incur other costs that have a primarily fixed nature, such as building a factory, production tools, marketing, and G&A, which need to be recovered through scaling the sale of products with positive contribution margins. Value from the physical good accrues to the company at the moment of a single transaction, where ownership of the good transfers from the seller to the customer.

How did SaaS change things?

Software products are different from physical goods. Software normally comes with sizable upfront development costs to create the software, while directly attributable costs to scale the software and deliver it to end users are virtually zero. Whether software is used by 10 or 1,000 users doesn't drastically change its cost structure (primarily the data infrastructure that supports operability of the software and customer support). So investors normally expect 70–90% contribution margins from SaaS products. The unit economics in the SaaS world becomes not the piece of software, but instead the human using that software.

The focus of these companies has turned from production to customer acquisition, retention, and upsell. Value from the unit economics of a SaaS product accrues to the company over the lifetime of the customer, not at the moment of sale. The key performance metrics in SaaS are tied primarily to the customer, not the product: for instance, LTV (lifetime value), CAC (customer acquisition cost), the LTV/CAC ratio, net revenue retention, magic number, etc. Even though the focus of the company became the customer rather than the product, companies are still pursuing scale to recover their fixed cost basis (tied to development and customer acquisition costs). SaaS products generate higher contribution margins than physical products, but they also come with higher uncertainty about the total value captured, as value accrues over time rather than in a single moment of sale. The strategy of companies becomes "acquire customers fast, retain as cheaply as possible, upsell as often as you can, and increase switching costs." Unit economics with low marginal costs incentivizes companies to aggressively pursue customer growth, making it the key topic of executives' attention.

What happened when Software AI entered the market?

Large language models arrived and reintroduced the thing SaaS had spent a few decades eliminating: a meaningful, recurring marginal cost. Every interaction with Claude or ChatGPT carries a variable cost driven by model consumption of GPU/CPU and energy. The benefits of the virtually zero marginal costs of a traditional SaaS product evaporated, as each query runs the model again, and the more information that query needs to account for (for instance, the history of prior communication), the higher the cost of a new query is. Agentic AI takes it even further, as each request to execute a task triggers a chain of reasoning, tool calls, information retrievals, and retries, causing the cost of a single session to multiply. Inference cost is frequently the largest variable line that scales with usage, presenting economic friction to increased utilization.

The unit economics in the Software AI world shifted from the customer to the query. Flat-rate subscriptions of a SaaS business break when heavy usage drives the cost per interaction substantially over the flat subscription fee, to the point where low usage doesn't offset it. That's why usage-based and hybrid pricing have returned, to price utilization closer to the cost to serve.

What also became important is whether the company selling a Software AI product owns its data processing infrastructure or rents it. Companies that are vertically integrated, owning the model and the data infra, carry the inference cost (data centers, energy) internally and have better opportunities and incentives to improve the efficiency of the model and the cost of data processing. However, building your own data centers is a huge deterrent for a software-native company, so many companies opt to specialize in the model and product, renting their data infrastructure from providers. This business model leaves companies with fewer opportunities to improve their contribution margins, as data infra providers charge per token with margins built into the pricing and have fewer incentives for optimization other than competition. More usage no longer means leverage; it means more cost. The discipline shifts from "acquire and retain cheaply" back to "serve profitably," a concept that is familiar to the physical product space.<

What is unique about economics of a Physical AI asset?

Physical AI assets include humanoids, autonomous vehicles, warehouse robots, drones, and other machines that perceive the world and act in it to fulfill a task. This class of products is commonly sold as a service: robotics-as-a-service, capacity-as-a-service, or labor-as-a-service. The unit economics that generates value becomes the work performed. The framing has shifted from "seat" in the traditional SaaS space to "work done"; value capture becomes more outcome-driven rather than customer-driven.

Economic value creation of a Physical AI asset has a hybrid nature; it doesn't fit perfectly into the profile of a traditional physical good or a SaaS product. The traditional model can't be applied because the asset is not sold, locking in margins at the point of sale; instead, it's deployed to deliver value to its consumer over time, so in that sense it resembles a SaaS product. However, the marginal cost of a Physical AI product is not virtually zero. This is an actual physical product that carries the cost of hardware, data, spare parts, installation, field operations, powering, repairing, human oversight, retiring, etc. This makes a Physical AI product unique: it has a recurring revenue stream normally associated with a SaaS product, paired with a capital-intensive production model associated with traditional physical products. The revenue from deployment of a Physical AI product now has to carry the burden of covering the upfront development costs, customer acquisition costs, installation/integration costs, cost of physical parts, cost to serve, etc. The value of the outcomes delivered to customers by deployment of a Physical AI asset has to be high enough to cover the costs and achieve the profitability expectations of investors now accustomed to SaaS margins and the adoption rates of Software AI.

The overall economic profile of a Physical AI asset is assessed by its ability to generate a stream of revenue over the full operating life cycle that covers upfront capital plus ongoing operating costs. This value converted into revenue is often tied to deliveries completed, items picked, or miles driven, but primarily to the cost of labor that the Physical AI asset displaced or the cost of the legacy machine it replaced. The core question becomes: can this unit deliver a unit of work at a fully loaded cost below the incumbent, at high enough utilization, to earn an acceptable return on the capital deployed? There are a few performance metrics that become critical to the Physical AI asset:

- Utilization. In a SaaS world, the revenue captured by the company from selling a "seat" on a subscription basis is the same regardless of how heavily the capabilities of the product are being used, or whether they are used at all. With a Physical AI asset the situation is different. If it's sold to the customer under a "$ per outcome" pricing structure, it only generates revenue when used. An idle unit might not be wasting ongoing operating costs, but it does waste the cost of capital used to develop it and the cost of a shortened lifespan, as a more advanced and capable machine is being developed to replace it. Physical AI unit economics is highly driven by uptime, operating hours, and throughput, similar to the logic that governs airlines, hotels, and equipment rentals. Utilization is a variable factor that has massive downstream consequences for unit economics and, hence, must become a key focus point for companies in this space.

- Scale. SaaS products don't get impacted by economies of scale as much as a Physical AI product does. Integration of SaaS products into the software universe of an enterprise is usually a tailored endeavor that creates meaningful revenue streams from software implementations and customizations. As the production volumes of Physical AI assets increase, Wright's Law kicks in, experience improves the efficiency of production methods, and the cost of components goes down with volume discounts. Physical AI assets experience a similar effect to what we observed with wind turbines, solar panels, and electric batteries. The more units are built, the cheaper each unit becomes. For companies building Physical AI products, the ability to scale, optimization of hardware architecture, and investment in production methods become critical factors impacting unit economics.

- Autonomy. Physical AI assets quite often operate with a human in the loop who supports the operational capabilities of the asset when the software lacks the capacity to handle a specific edge-case scenario of reality. Think of Waymo's teleoperation, where a remote assistance operator gets involved in moving a vehicle out of a situation that the software can't handle. The labor costs of human supervision go into operating costs, dragging margins down. As software capabilities improve, allowing for more autonomous operation, unit margins improve. Quite often, improving the software model requires data coming from deployment of the assets; this is how scale becomes another contributor to the unit economics, by creating a data flywheel that feeds data into model training, which in turn increases the autonomy of the Physical AI asset.

- Upgrades. When companies improve the cost per unit of a traditional product through volume discounts, production methods, etc., those improvements do not go backward to improve the cost of units that have already been produced. There is no backward cost improvement. Physical AI assets, however, do enjoy the benefits of software improvements, provided that the hardware of the units allows for it. When the capabilities of the software improve, those improvements can be sent wirelessly to the entire fleet of Physical AI assets, raising the unit economics of all of them. Scale and a proper data flywheel enhance their importance through their ability to impact the unit economics of the entire fleet of Physical AI assets. What also becomes important is designing hardware in a way that accounts for future software improvements, not just current software capabilities at the moment of unit production to fulfill current customer requirements.

- Ownership and financing. With traditional products, ownership is usually transferred to the customer at the moment of sale, enabling the customer to capture all benefits from using the product. On some occasions, the seller finances the product, either itself or through a participating financial institution, but value still accrues to the customer after the transaction closes, as the product is being used. The seller captures the value at the moment of sale. With a SaaS product, both customers and sellers capture the value over time. Customers use the product to achieve certain outcomes over time and pay a subscription fee to the seller; in other words, financing is in a sense included in the subscription fee structure. The ownership and financing structure of a Physical AI product can fall anywhere between the two models. Let's consider two scenarios:

1. The seller retains ownership of the assets, selling their utilization to the customer on an "as-a-service" basis. This resembles a SaaS model, where the seller accrues value over the useful life of the asset. However, if the revenue is tied to the outcome and not to time periods, the revenue stream is not as predictable as that of a SaaS product. What complicates the matter even further for the seller is that the seller has less control over utilization of the asset by the customer, so while the seller retains ownership of the asset and all the risks and costs associated with that, the customer is given a lot of control over the revenue stream that affects the unit economics of the asset on the seller's books. The seller carries the upfront development costs, the cost of improving capabilities, the cost to serve, and capital costs, but doesn't fully control the revenue stream needed to cover those costs, and also needs to carry the cost of incentivizing increased utilization by the customer.

2. The seller transfers ownership of the Physical AI asset to the customer. In this case, the customer has incentives to utilize the product to the maximum of its abilities, as they capture full value over its useful life. However, if the seller accrues its value at the moment of the transaction, it loses the incentive to invest in improving asset capabilities. Those incentives and performance requirements become hours of contractual negotiations between the parties, turn into pages of legal language, extend time to adoption, hinder the pace of scaling, create the need for compliance monitoring, and create all sorts of friction.

3. A hybrid model, where either the seller or the customer takes on the burden of ownership and associated risks and gets compensated for those risks accordingly. The business model and pricing architecture are designed in a manner that creates alignment of incentives between the seller and the customer and allows for distribution of the value from deployment of a Physical AI asset between the parties, where each party gets rewarded fairly for the contributions they make and the risks they carry. This structure should be carefully constructed by finance that understands the complexities of economic value creation and distribution in the Physical AI space.

The economics of a Physical AI asset depend on the cost to develop, the cost to build, the cost to deploy, the cost to serve, the cost to scale, and the cost to improve (capabilities, autonomy, utilization). They account for the full life cycle of the product, similar to how project finance for real estate projects works. The usual metrics of LTV/CAC or gross margin at the point of sale lose their relevance in favor of asset-level NPV, IRR, cost per unit of output, payback period, utilization rate, gross operating margins, etc.

What does this mean?

Physical AI doesn't introduce a new kind of unit economics so much as it inherits every kind that came before. It carries the capital intensity of physical goods, the recurring value capture of SaaS, and the comeback of the marginal cost of Software AI. The businesses that win will be the ones that create business models addressing all of these forces at once. A Physical AI product has complex unit economics that take into account the entire lifecycle of the product and the interdependencies between the various factors that impact a company's ability to capture the value the product creates. There's a reason why getting the unit economics of a Physical AI product right from the onset matters more than it did in software. When a SaaS company gets its unit economics wrong, it burns cash, learns, and can often pivot its way out. When a Physical AI company gets its unit economics wrong, the mistake is already sunk into hardware architecture, production facilities, field operations infrastructure, price per outcome expectations, etc. That's why a finance function fluent in the economics of innovation isn't a back-office hire in this space, it must be part of the founding team. The cost of getting unit economics wrong substantially outweighs the cost of bringing finance into the conversation early enough to do it right.

June 2026

← Insights

AI Can Build the Model. It Can't Broker the Trust Between Capital and Organization. That’s Strategic Finance’s Job.

Every week brings a new headline about AI taking over finance jobs. I spent the majority of my career in strategic finance. Like many people in intelligence-heavy fields, I felt uneasy with the narrative of AI commoditizing intelligence. However, after diving deeper into transformer architecture, learning more about the foundational principles of AI, and extensively using its capabilities myself, my view on AI's role in finance has shifted. AI has commoditized language, but not the financial intelligence that is at the core of strategic finance's value proposition. Judgment, relationships, accountability, creativity, seeing a bigger picture — these are uniquely human qualities that make up the bulk of the strategic finance skillset. I see AI as a great collaborator for strategic finance, not its replacement.

Before we dive into use cases for AI in strategic finance, let me first explain how I see the role of strategic finance in the organization.

What is strategic finance?

So many times in my career people used strategic finance and FP&A (Financial Planning & Analysis) interchangeably. I can see how for someone outside of finance, those two sound the same. The scope is presumed to include budgeting, forecasting, management reporting, and ad-hoc analytics. And that’s all true. Those are primary activities covered by the FP&A function.

The scope of strategic finance, however, is much broader than that of pure FP&A. Strategic finance is primarily concerned with

- shaping the strategic direction with an overarching goal to achieve maximum return on capital deployed within the capitalistic system,

- aligning functional activities within the organization around a single vision of the company’s financial future,

- developing and leveraging core organizational competencies to generate value demanded by the market,

- structuring value capture mechanisms, including business model and pricing architecture, best suited for the nature of the product and market environment,

- utilizing market forces to capture and retain value in a capital-efficient manner, and

- investing in a competitive moat to secure market positioning and future prospects.

Strategic finance is a broker of trust between capital and the organization that deploys it. The role of strategic finance is very similar to the role of government that acts as a broker of trust between people paying taxes and government’s bureaucratic machine that deploys those taxes to supply citizens with food, shelter, safety, and social structure that provides opportunities for a better life. A similar line of thinking applies to the venture capital industry. VC firms exist to broker trust between capital owners (investors in VC funds) and entrepreneurial opportunities with high potential of outsized economic returns. Investment banks broker trust between institutional and retail capital and the public market. Governments, venture capital, investment banks, and strategic finance serve a similar purpose just on a different scale. Strategic finance is embedded inside companies with a mandate to ensure that the company deploys capital in a manner that has the highest chance of generating maximum return.

How can AI help?

This is where I see AI adding value to strategic finance.

- Scaffolding of the financial models.

I always built my own financial models from scratch. There is a certain level of creativity involved in building the structure, assessing the need for its depth and complexity, investigating the key assumptions, and envisioning the conversations the model will facilitate. I never needed to use AI to help me with that. Recently, as I investigated the economic potential of various frontier technologies, I tested the model-building capabilities of AI tools, asking them to build financial models using publicly available information. They did build the models quickly, burning almost all of my token allowance, but they did a good job creating draft versions. File structure, assumptions, three-statement linkages, scenarios, sensitivity architecture, Monte Carlo setup, etc. A few errors here and there, but generally speaking, it produced a solid foundation for financial models that can be made useful via refinements that reflect nuances of the specific business and its environment, use case, target audience, granularity, and support for the key assumptions driving the financial outcomes.

- Research assistant.

AI chatbots have replaced Google for me as the go-to source of information for intelligence gathering exercises. AI is the most powerful technology I’ve seen for analyzing massive amounts of data in a short period of time to extract relevant insights and package them in the format convenient for consumption. I use it for analysis of market research, industry trends, government regulations, tailwinds and headwinds, competitive dynamics, and assumptions that go into financial models. It’s a great tool for synthesizing market intel, drafting narratives, creating storylines, helping with investor Q&A prep, and condensing content into bite-size pieces. However, there is human judgment involved in making use of the intelligence gathered by AI. There is also a very strong need to verify sources and the accuracy of interpretation of information by AI. Strategic finance influences multi-billion-dollar decisions; the tolerance level for hallucinations or misreading the data is much lower than in most other use cases of AI. I recommend tracing every model assumption and critical input back to their sources that can survive legal and investor scrutiny. And AI cannot account for all relevant information and intuition that it doesn’t have access to.

For instance, take a frontier hardware company — a robotics startup. Building a credible financial model means making assumptions about things that don't have clean historical data yet: component cost curves as production scales, rate of improvement of human-in-the-loop intervention, supply chain dependencies, useful life of components and rate of repairs, customer adoption rates, willingness to pay, revenue structure, milestones unlocking new revenue streams. A human analyst could spend a week pulling this together from different sources scattered across the internet. AI can compress that into an afternoon, surfacing the comparable cost curves, flagging assumptions that lack public data and need a judgment call, and drafting a first-pass sensitivity range for the assumption. What it can't do is decide whether the rate-of-repairs assumption is aggressive or conservative given what the VP of Engineering said in last week's product review, whether the assumption on customer willingness to pay on a recurring basis is realistic based on a conversation that the VP of Business Development had at a conference last month, or whether the investors will find the path to breakeven credible enough to fund the next raise. The judgment calls — what public source to rely on, what historical precedent to use as a reference point, what logic to deploy where data doesn't exist, which market forces shaping the assumptions deserve attention, what feedback from internal stakeholders to incorporate or dismiss — those are still the strategic finance leader's to make.

- Automation of routine tasks.

Agentic AI is useful for automating routine, repetitive tasks of high value. For instance, combining data from different sources to automate management reporting, variance analysis, and re-forecasting. AI can be used to automatically update KPI dashboards; it can analyze market signals, signals from internal communication channels, and competitors’ moves to raise alarms that a specific scenario is unfolding. It can even automatically suggest strategic responses to external and internal events to trigger further discussions amongst stakeholders. This is the area where agentic AI can be incredibly useful to strategic finance, increasing the velocity and depth of relevant action-focused conversations.

- Always on.

It has no ego, no ambitions for career growth, always happy to receive feedback and criticism, and has infinite patience to redo the job as many times as needed, of course, as long as you have your token allowance.

What are humans better at?

This is where I see strategic finance requiring human touch.

- Making financial models useful.

Financial models do not exist in a vacuum. They are not created to impress investors during fundraising and put on a shelf. They are created to communicate current financial reality, present versions of the future out of the many possibilities that could unfold, and facilitate conversations. Financial models serve multiple purposes:

- inform the leadership of the company’s financial performance,

- identify key operational drivers that impact that financial performance,

- interpret relationships between different operational drivers,

- keep a placeholder for key assumptions for the values of those operational drivers,

- suggest behavior patterns for how those values will change over time and how that will impact the company’s financial results,

- demonstrate different versions of reality in response to the behavior of different operational drivers (for instance, via Monte Carlo simulation or scenario planning),

- present trade-offs and financial outcomes of different decisions,

- become an anchor of conversations amongst the key stakeholders to align on the operational areas with the highest economic value to the company.

The true value of a financial model comes not so much from the mechanics and formulas, but from the hypothesis behind its assumptions, logic incorporated into inputs-outputs relations, and usefulness in discussions with stakeholders. Model assumptions can be backed by historical numbers, tied to a credible source, hypothesized to reflect market forces and competitive dynamics, or simply act as the targets for what the respective stakeholders should strive them to become. The model is not meant to be a static tool; it’s meant to be a dynamic asset that facilitates decision-making. And that requires judgment.

- Applying judgment.

Judgment is not the same as reasoning. AI has learned to reason, but its reasoning is based strictly on data that it had access to when making that reasoning. What it cannot account for is the multi-faceted nature of the organizational communication web, differences in how different stakeholders consume information, who needs to be brought in to create initial inputs, and who prefers to have their inputs incorporated into the pre-final draft of the model, how to tailor information and narrative behind it to a specific audience, what stakeholders actually think but don’t want on the record, what they shared in a water-cooler conversation about the ground truth of reality. Organizational behavior and the realm of ideas are still intangible objects that exist in the collective consciousness of the employees. AI simply cannot have access to those (yet) and cannot apply judgment about how to leverage the financial models to drive productive conversations and influence decision-making.

- Influencing decision-making.

Strategic finance is primarily a contact sport: getting the GM to agree to and own their budget, negotiating pricing guidelines with business development, narrating the financial story to the investors, knowing whose number to trust and whose to discount, understanding whose support to obtain before bringing up a strategic discussion to the key decision-makers. The success of strategic finance in guiding or influencing decisions in the company is heavily dependent on its relationships with stakeholders, ability to read the room, skillful balancing of when to push back and when to let go, historical knowledge of what worked and didn’t work in the past, and intuition that is built over time. AI lives in the soft world, while decisions are made in the physical realm, at least for now. I know how to identify relevant intelligence, how to read competitive dynamics, when to treat a signal as noise and when to treat it as an indication to take action, when to raise red flags, when to pivot, when to stop the cash bleeding, when to throw more capital into the initiative, when to reallocate resources, when and how to initiate decision-making conversations. AI is not capable of any of that.

There is an argument that AI is capable of agency to make decisions and it’s up to organizations to leverage that capability. I don’t see that being the case. If AI is allowed to execute its agency in the strategic finance field and make decisions on deployment of capital with strategic and financial consequences on its own, then we don’t need the C-suite, we don’t need the Board, we don’t need venture capital, we don’t need investment bankers, we don’t need government. As I mentioned before, strategic finance is a broker of trust between capital and the organization with the mandate to ensure its efficient deployment. If capital deployment decisions are outsourced to AI, then the entire financial system must be managed by AI to deliver optimal outcomes. Perhaps that's the future we'll build — one where AI manages the entire economy. However, for now, the entire financial system is an architecture of accountable trust, and you can't delegate trust to something that can't be held accountable.

- Taking ownership and accountability.

The head of strategic finance signs their name to the forecast, acknowledging the reasonableness of key assumptions and the high likelihood of reality unfolding according to the projections based on the best available information at the moment and the best assessment of the commitments from the key stakeholders with the power to shape that reality. They carry real consequences that impact their promotion, pay, and power within the organization. AI doesn’t carry any accountability for its advice or its consequences. It can't be fired, can't lose credibility with investors, and doesn't sign its name to a forecast or a decision.

I have been in numerous conversations with lawyers and investors defending the company’s unit economics, financial projections, and key assumptions driving those projections. I cannot imagine lawyers and investors having those conversations with an AI chatbot and relying on that conversation to sign off on numbers that go to the SEC or to invest in a fundraising round. Strategic finance is capable of painting a bigger picture, emphasizing points that are relevant to the conversation counterpart, applying different approaches to instill confidence in numbers based on the circumstances, and incorporating feedback or valid points raised in the conversation into the next iteration of the financial model, often in the same meeting.

- Reproducing thinking patterns.

Reproducibility of the outcomes of AI models is one of the key sources of friction in adoption of AI in finance. Stakeholders of finance expect reliability of data, judgment, logic, and advice. One of the criteria that establish the reliability level is: “can you do it again?” While AI can replicate creation of assets — KPI dashboards, board financial update slides, budget vs. actual review decks, and other management reports — it has a hard time replicating the reasoning and thinking patterns. I know the support behind each financial number communicated to any stakeholder in any situation and I can explain how it came to life. AI can change its reasoning on its own, taking different reasoning paths to produce results. When strategic finance engages in alignment of stakeholders around capital allocation decisions, that variability is not helpful. Strategic finance needs to have a single foundational narrative and tailor its delivery to the stakeholders’ roles. Reproducibility of thoughts and logic is still a human quality.

- Leveraging creativity.

AI-generated language content is not original.The models are built with the information that already exists. Data that feeds the algorithm of an AI model has already been created. Since AI models are trained on the data available to all AI model producers, they are generally producing homogenized commoditized outputs — more of the same. Humans, however, have unique abilities to find creative solutions out of pure imagination, finding inspiration in sources that we don’t fully understand. Beginner’s mind, a concept defined in Zen Buddhism as “an attitude of openness, curiosity, eagerness, and a total lack of preconceptions when approaching a subject”, is a uniquely human quality. Strategic finance often works with incomplete information in ambiguous settings to imagine what the future might and should look like and this often involves finding novel solutions and trying out new approaches to problem-solving. The imagination and creative force of the human mind are among the key skills deployed by strategic finance leaders to tackle situations not reflected in the existing data. The human mind can tap into a much vaster pool of data, feelings, imagination, and intuition, something that AI simply cannot do due to the limitations of its training set and underlying architecture.

So where does that leave strategic finance in the age of AI?

AI leaves strategic finance leaders with a powerful set of tools enabling focus on higher-value-added activities. When AI capabilities are used well, AI can deliver the output of a small team of analysts: it can scaffold the models, synthesize the market intel, draft the commentary, and automate the routine so that human attention flows to what actually moves the needle. I see the division of labor between humans and AI in strategic finance working as follows:

- AI drafts the models; the human refines them and owns the assumptions.

- AI surfaces the market signals; the human decides which signals to treat as noise vs. which signals warrant action.

- AI produces the analysis; the human updates the forecast, develops recommendations, orchestrates decision-making, and carries the consequences.

The narrative of AI taking over finance is overblown. In the 1980s, spreadsheets were supposed to eliminate accountants. Instead, they eliminated menial bookkeeping tasks while multiplying the productivity of the profession. The same is happening in the field of strategic finance. The tangible products of strategic finance — the model, the presentation, the dashboard — are not the point of the function. The point of strategic finance is to manage trust between capital and an organization guiding the deployment of that capital in a financially sound manner while accounting for the context in which the organization operates. Judgment, relationships, accountability, creativity, seeing a bigger picture — all of the work that happens to turn the model into the decision is what strategic finance is owning. With the cost of capital rising, organizations need to apply a more rigorous financial discipline to the deployment of capital. AI can build the model. It can't broker the trust. And as long as capital is deployed by humans, it’s the humans with skin in the game who need to stand behind the numbers. Powered by AI, strategic finance can stand on firmer ground, investing more time in the areas that matter to the financial success of organizations.

July 2026

← Insights

Robots. When Are They Going to Make My Breakfast?

A few weeks ago, I attended an event at Stanford with the topic of Humanity and AGI. There were a few companies that used the event as a marketing channel to demonstrate the capabilities of the humanoids they've built and are selling to consumers, some priced at $6,000 a piece. Two humanoids dancing on stage fell off it; one shattered the presenters' TV. When another humanoid almost fell on a person while taking the stairs to get off the stage, I wondered how close we really are to having a robot making an omelet for breakfast. With all the advancements in the AI space, why is it so hard to make a humanoid that can be invited to our homes to give us relief from activities we crave to outsource to robots the most, the house chores? As a strategic finance expert in frontier tech, I saw the performance failures of humanoids at the Stanford event as clear signals that developing robots that can actually be useful must be a very expensive endeavor. I wondered how strategic finance can get involved to help the companies in the robotics space to manage engineering challenges in a financially sound manner.

My investigative journey led me to a key conclusion that the humanoid industry has a structural bias towards spending capital where physics makes data cheap (locomotion, demos) rather than where the economic value actually lives (manipulation, hand dexterity), and closing the gap between demo-ready and commercially-ready humanoid is as much capital-allocation problem as it is an engineering one. Let me take you through the journey.

Robot

Aristotle said in his Politics: “If every tool, when ordered, or even of its own accord, could do the work that befits it... then there would be no need either of apprentices for the master workers or of slaves for the lords.” The idea of having robots supporting the advancement of humanity existed in our imagination since at least 350 BC.

Isaac Asimov defined a robot in his fictional universe as a computerized, mobile machine equipped with a "positronic brain" capable of performing complex tasks and governed strictly by the Three Laws of Robotics:

First Law: A robot may not injure a human being or, through inaction, allow a human being to come to harm.

Second Law: A robot must obey orders given by human beings, except where that conflicts with the First Law.

Third Law: A robot must protect its own existence, except where that conflicts with the First or Second Law.

Practically, I see two versions of robots:

1. Programmed Robot — a machine equipped with mechanical and electronic parts, which can accomplish a prescribed task in a physical world without the need for understanding it.

2. Intelligent Robot — a machine that requires an understanding of the world to accomplish a prescribed task in a physical world using its mechanical and electronic parts.

In both versions, Robot is a physical, tangible object. It is instructed to achieve a task, transforming one state of the physical world into another state. It has hardware and software that govern its operation. It has mechanical parts (body, legs or wheels, arms, hands, pickers, fingers) that interact with the physical environment and physical objects. It has electronic parts that make the mechanical parts move to manipulate the state of reality. However, a Programmed Robot, like the industrial robotic arms that assemble cars, does not understand what piece of a car it attaches to what piece of a car and why; it relies on humans and the operational process design to ensure that the right pieces are put together. Robots that run through Amazon warehouses to move packages have dedicated lanes and don’t need to understand the layout, just the lanes they are assigned to operate in. Programmed Robots don’t need to make judgment calls; they operate with a clear set of instructions in a structure environment. If they encounter situations not covered by their code, they simply stop execution of the task until a human assesses the situation and restarts the operation. Intelligent Robots, like autonomous cars, make judgment calls all the time. They need to understand which part of the road is safe to drive on, what road objects to avoid or fine to drive over, when to yield to another car or a human, when to squeeze between the cars changing lanes. They constantly analyze the physical environment, understand surrounding objects, make predictions, create alternatives for behavior, assess the cost function of each alternative, and make a call on what action to take with safety being the primary objective. For this article, I want to focus on a humanoid as a type of Intelligent Robot that has a lot of promise but also a lot of complications on the path to commercialization.

We entered a new stage of evolution with the large language models (LLMs) powering functionality of ChatGPT, Claude, Replit, and other AI tools. However, understanding the unstructured physical world is extremely complicated, more so than understanding the language world. Per Fei-Fei Li’s essay, A Functional Taxonomy of World Models: “Language models have given machines an extraordinary command of concepts, vocabulary, and reasoning, but the physical world, virtual or real, runs on a different substrate. Where language models learn the statistical structure of text, world models learn the statistical structure of space and time: how light falls on a surface, how a garden looks from an angle no camera has captured, how objects respond to force and follow the laws of physics.” Learning to understand the physical world requires a new codified version of the world, a different kind of data than the one that fed LLMs, a vast amount of that data, and maybe even a different kind of learning. Let’s break the complexity into pieces.

Hardware

Humanoid is not software, it’s a physical object navigating through a physical space and interacting with physical objects. Programmed Robots have been deployed in factories for decades, since the Industrial Revolution. But Intelligent Robots is a fairly new phenomenon; they haven’t gone mainstream yet. Deploying Intelligent Robots to the environment designed for humans without radically changing is the holy grail, it will unlock trillions of dollars of value. So the robotics industry and Venture Capital have dedicated a sizable amount of capital to developing humanoids that replicate the functionality of a human body. That would make their adoption easier as it won’t require a major overhaul of the physical infrastructure to support their operation in a world designed for humans. However, that path is not without its challenges:

- Movement.

The human body is a piece of art. Da Vinci’s Vitruvian Man is one of the most recognizable and all-time iconic images of Western civilization. The human body has ~350 joints enabling over 200 skeletal degrees of freedom (DoF), with approximately 80 utilized for everyday whole-body motion. The human hand itself is a nature’s miracle with 27 degrees of freedom, capable of performing virtually unlimited sets of tasks that created the world as we know it. Replicating a fully human body and human hand in a humanoid is extremely challenging. Humanoid robots typically are built with 30 to 50+ degrees of freedom. A standard breakdown allocates roughly 12 to 14 DoF for the legs, 2 to 4 for the torso and neck, 6 to 7 per arm, and up to 20+ if using multi-jointed, dexterous hands. Each joint is a mechanical part with sensors that must generate data for training, execute commands when in production mode, withstand extensive hours of use, and be easily serviceable. Humanoids developed for demos might not need to handle the complexity that comes with dexterity. However, for commercial applications, dexterity is one of the key drivers for customers’ willingness to pay and adoption rates. Requirements for dexterity of the body and arms come with major hardware challenges, data requirements, and heavy price tags for the developer and the eventual consumer.

- Power.

Humanoids need to be able to operate without being constantly plugged in. Most humanoid platforms on the market can operate for two to four hours on a single charge. That might be okay for personal use and a task to clean a house, but not sufficient for deployment in industrial settings. And that challenge can’t be just solved by choosing a bigger battery. A humanoid must carry its energy source inside a human-shaped body and balance that mass over two feet. A wheeled design might address that challenge by storing a bigger battery in the base of the humanoid and reducing power demand for balancing the robot. But wheeled humanoids substantially reduce the scope of practical applications. Bipedal, or legged, humanoids, like Tesla Optimus and Figure 03, have a much broader set of use cases making them significantly more commercially attractive. But bipedal design choice creates limitations for the size of a battery and where that battery can be placed without wrecking the robot’s center of gravity. Advanced battery technologies like solid-state cells, silicon anodes, and high-energy lithium-sulfur chemistries could solve the energy problems by vastly increasing energy density, but those improvements must outpace the additional energy demands created by rising complexity.

- Technical limitations.

Use cases for humanoids include deployments for dirty, dangerous, and demeaning jobs (3D). Think of the following examples: dirty - handling unsanitary or toxic materials, dangerous - repairing outside of the space station, demeaning - monotonous sorting cycles that reduce humans to automation machines. There is a high demand to deploy humanoid robots in environments involving extreme physical risk or extreme conditions. Making standard robot joints and wiring harnesses durable enough to survive these conditions adds weight, cost, and power demand. Harsh environments degrade mechanical joints quickly. Dust, water, and extreme heat ruin sensitive electronics. Walking on uneven, slippery, or shifting surfaces and handling steep stairs and tight spaces is very difficult for humanoids. Imagine what happens if a humanoid carrying hazardous material loses power or one of the knee joint malfunctions, making the robot fall due to gravity; that can cause a secondary accident with catastrophic consequences. Underground or heavily shielded facilities can block wireless signals, making real-time control and monitoring impossible. 3D tasks might require special tools that place additional demands on the dexterity of the robotic hand, requiring dozens of miniature actuators, sensors, and cables packed into a tiny space. This intricate wiring is highly prone to failing under rough conditions.

These technical challenges are very expensive to solve. They are highly dependent on the use case, cost of engineering roadmap to solve them, and customer willingness to pay. Enhancing hand dexterity, solving for durability and withstanding harsh conditions, addressing power limitations, and expanding versatility of deployment use cases impacts BOM, maintenance costs, warranty, replacement cycles, and cost-to-serve, i.e. TCO, which is the number the buyer actually cares about. 2-4 hour battery life is not an inconvenience; it's a utilization ceiling that drags down ROI and caps the value a buyer gets from robot’s deployment. In my own experience, I witnessed the consequences of engineering roadmaps designed without finance inputs, setting engineering objectives around demonstration of capabilities at the expense of investing in the actual capabilities, failing to account for full cost impact of hardware architecture in the product pricing, and scaling with negative unit economics. Those are expensive decisions that can be avoided if finance is brought to the table to partner with the C-team ahead of time.

World Model

We don’t know with 100% certainty what the physical world is truly made of (particles, energy), or even whether fundamental reality is not material, but mental or experiential, and what we call matter is how that reality appears. But we do know for sure that physical reality is not made of words. We interact with the world through our senses. We feel it through vision, hearing, smell, taste, and touch. We have spatial awareness to navigate through the physical world without stumbling over objects on our way. We recognize objects and their functional purpose even before we assign a word to that object. We understand how to interact with objects in a manner that accomplishes a task created in our mind. We have the capability to predict the outcome of our interactions with the physical environment and the objects in it before we execute the action. By observing how the world and the object respond to our action, we build a more comprehensive understanding of the range of possibilities for our world interaction model.

And that’s what I have in mind when I say “world model”. It’s not the simulated version of the world like the one you see in the video game. It’s a dynamic interaction model between physical action and the state of an environment. A transformer LLM is trained on next-token prediction: given a sequence of tokens, estimate the probability distribution over the next one. A world model predicts the next state of an environment, conditioned on an action. World models for humanoid robots are generative neural networks that encode an environmental state, predict future physical consequences conditioned on a robot's actions, and simulate physics in a latent space. LLM is text-in, text-out. World Model is action-in, future-world-out.

There are many competing schools of thought on how a robot should model the world, or whether it needs an explicit world model at all. There is a Latent Dynamics Model, which predicts how an environment evolves inside a compact, abstract mathematical space rather than generating full pixel-by-pixel videos. The most prominent example is Meta's V-JEPA 2. There is a Generative Video Prediction Model, which creates realistic, high-definition future video frames of a scene given the robot’s current camera view and an intended movement. Nvidia Cosmos, DeepMind's Genie are amongst the most instructive production examples. There is a World-Action Model, which is not really a separate model, but a fusion approach that blends video generation and continuous motor-control policies into a single architecture, mapping both future visual outcomes and precise joint movements. There is also a camp that argues that you don't need an explicit world model at all. Figure's Helix, Physical Intelligence's π₀ are direct vision-language-action imitation: map observation plus instruction straight to action with no future-prediction step. Which approach is the most efficient and future-proof, only time will tell.

Regardless of what world model architecture is used by the humanoid developer, model training needs data, and data for training intelligent robots is not cheap: text is nearly free and effectively infinite, while physical-interaction data has to be generated through simulation, teleoperation, human video demonstration, or real robot time, all of which are slow and expensive. The most valuable stream, the action, is the one the internet can't provide. There are endless videos of people doing things online, but with no record of the motor commands that produced the motion, making those videos less useful for modeling interaction. It’s not just the data economics that is a bottleneck, it’s also data complexity, variability, duration, and interrelations that real-world deployment demands. Data economics is a critical area that requires the involvement of strategic finance. What data flywheel to build, what data to collect, what data to manufacture, what functionality that data should enable, what customers are going to pay for that functionality, and how that impacts hardware architecture and margins of commercial units — these are finance questions that need answers before data consumes the capital provided by investors.

Data

Let’s dive deeper into data.

The most advanced consumer-facing robotics product, in my opinion, is a robotaxi. Waymo has been taking over the streets of major cities, demonstrating that it is possible for a machine to fulfill a task operating in a physical world sharing it with regular people. By summer of 2024, when Waymo opened its robotaxi to all San Francisco users, Waymo had driven about 20 million fully autonomous miles on public roads vs. tens of billions in simulation. This gap is exactly the economics of physical data. And that was for modeling behavior in a constrained environment — roads designed for vehicles. The surface where autonomous cars can operate is clearly marked (mostly), there are rules of behavior for all participants, there are road signs that apply to all participants, the primary goal is to avoid collision with other vehicles and objects on the road (people, dogs, stationary and moving obstacles). We designed roads for vehicles and placed intelligent robots inside that structured physical environment. Now, if we want intelligent robots to operate in space shared with humans or even come inside our house, that’s a whole different game. That environment is unstructured with more ambiguous rules of behavior. There are a lot more objects that interfere with achieving the task and a lot of variability of those objects. The cost function that drives choice between alternatives is unclear. The interaction model with the world is more physical (touching, manipulating objects rather than avoiding them, as autonomous vehicles do).

Deployment of intelligent robots in the real world has a vastly more complex functional task than the autonomous vehicles. The goal of an intelligent robot is not to avoid the objects, but to interact with and transform them from one state to another. There are many nuances to consider, a few I can think of:

1. What kind of data to collect for training? Many companies collect video data and physical motion data to build connections between mechanical interaction with the objects and visual representation of different states before and after the interaction occurred. This is similar to how a child experiences the world, seeing, interacting, and understanding the connection between the two.

2. What point of view to take? We, humans, experience the world from the ego-centric point of view; we see the world with our eyes and place ourselves at the center of experiencing the world around us. So many companies collect data from the same ego-centric point of view of a robot connecting video feed from the camera with the physical motion of the arms controlled by a human through teleoperation. The idea is to collect data from the ego-centric point of view of the robot connecting its perception with the mechanical motions of the arms/legs. However, humans can also imagine the world as a 3D space outside of their own selves that helps us with predictions of how the world will change in response to our actions. We can create a 3D model of every object we interact with to assist us with achieving the purpose of that interaction.

3. What / how many sensors to build reliance on? Robot needs to understand its surroundings and its placement in those surroundings. It also needs to understand the nature of the objects it is expected to interact with to accomplish the tasks and identify those objects in its environment. Humans have a very good depth perception and very sophisticated body that supplies the brain with sensory experience data. We understand how gentle we need to be with the object or how much power to use for handling a heavy object. We can even manipulate the objects without seeing them, just relying on the sensory inputs. That poses questions about whether we should assign different purposes for the visual perception sensors and tactile sensors? Robot needs to understand how to interact with the objects to accomplish the task without breaking them. That involves building a connection between observation and action, how to manage mechanical parts (hands, joints, grippers, force, tactile sensing) of the robot to control interaction with the objects. Collecting sensory data is the problem that many companies are trying to solve, but that data is dependent on the hardware itself, so if hardware changes, the sensory data that controls the movement of that hardware is likely to change as well. We’ve seen videos of robots breaking eggs and making omelets. But can that robot also pick up salt shakers that take different forms in our kitchens and put the right amount of salt into the egg mixture? Can it measure and put the flour into the mixture to make muffins? That’s an extremely complicated multi-faceted data problem from both visual understanding and sensory experience perspectives that is hard to solve without a meaningful capital infusion.

4. How to convert world interaction data into training data? How to represent physical motion data? What kind of labeling approach to apply to perception, motion, and action data? Children learn to interact with objects before they can label them; they learn to understand the functional purpose of different objects without the need for language. So is it helpful to go through the pain of labeling objects?

5. When to use real data and when to use simulation? Simulation generates trajectories at near-zero marginal cost with perfect labels and supports reinforcement learning where trial-and-error is free. The cost is the sim-to-real gap — contact physics, deformables, and friction are exactly where simulators are weakest and exactly what manipulation depends on in the real world deployment.

6. How to leverage data from the deployment of the robot itself? Establishing a data flywheel to collect data from the deployment of robots that feeds into retraining of the model allowing for continuous improvement is a great idea that came from the software world. However, expectations for capabilities and safety metrics of deploying hardware in environments where they interact with humans are much higher. The robot is expected to execute its functional purpose in a safe manner before it can be commercially deployed, so leveraging deployment data might prove to be difficult for certain robot use cases.

Training a humanoid takes a massive amount of data. And not just one dataset, like text that is used for training LLMs, but several distinct data streams, each answering a different question about the robot's understanding of the world around it and supporting execution of the assigned task. I group these data streams by their purpose:

Vision - “where I am and what's in front of me?”.

This understanding comes from cameras on the head or torso to read the scene, and often from cameras mounted on each wrist for a close-up view of whatever the hands are manipulating. The wrist views matter because they see the contact up close and don't get blocked by the robot's own body. In some situations, it might be helpful to add LiDAR sensors that sharpen depth perception and spatial understanding.

Proprioception - “what shape I am in right now?”

The core signals are joint positions (the angle of every joint) and joint velocities (how fast each is changing), read from encoders on each motor. From those angles plus the known lengths of the links, forward kinematics computes where every part of the body is in space, where the hand is, where the elbow is, how the torso sits relative to the hips. This is entirely internal: it tells the robot's pose without any reference to the outside world or to gravity. A robot lying on its back and a robot standing upright can have identical proprioceptive readings if their joints are bent the same way.

Balance and inertial state - “which way is down, am I tipping, and am I about to fall?"

The anchor sensor is the IMU (inertial measurement unit), which measures linear acceleration and angular velocity — the same principle as a human’s inner ear. From the IMU, plus the proprioceptive pose, plus the foot force readings, the system estimates which direction is truly vertical, where the center of mass sits, and where the center of pressure under the feet falls relative to the support area. This is the only one of these streams that is about the robot's body as a single object oriented in the world, not about any one joint or contact point. Balance is also what couples everything together. When a humanoid reaches for a cup, that movement shifts its center of mass, which forces the legs to compensate, which changes the torso pose, which moves the cameras and the arm's base frame. Manipulation and locomotion aren't separable in a humanoid, because they share one body, one balance problem, and one power budget.

Force and torque - “what is pushing on me, and how hard?"

These come from dedicated load cells, usually six-axis sensors at the wrists and ankles, plus torque estimates at the joints. Where proprioception is not designed to understand contact with the environment, force and torque are entirely about the boundary between the robot and everything it touches. When the gripper presses a button, when a foot lands on the floor, when a hand lifts a heavy box versus an empty one — proprioception can't tell those apart, because the joint angles might be identical, but the force sensors read completely different loads. This is the stream that helps with questions: "how firmly I am grasping," "did I make contact yet," "is this heavier than expected," "is my foot actually bearing weight." It's the sense of effort, not the robot’s position.

Action - “what I am trying to achieve?”

Everything above is input data describing the robot's situation. The action is the output: the actual commands issued to the mechanical parts to execute the task. This is the answer the model is being trained to produce. It's the single most valuable stream and the single most expensive, and it only makes sense paired with all the observations around it — time-synchronized egocentric video, joint positions and velocities, torques, and tactile readings, alongside the command that was actually sent. Actions are collected through teleoperation, motion-capture suits, egocentric human video, or simulation. Natural-language task instructions are paired with the specific physical episode they describe. Instruction to "pick up the mug" is attached to the exact trajectory of movements that did it. This is what lets the robot connect a goal to a physical behavior.

If the above wasn’t enough to signal the complexity of the AI model training for humanoids, I want to mention a few data challenges:

- Time synchronization.

There is a difference in how quickly a model can process incoming information from perception and motion sensors and how quickly motion sensors need to react to a change in the environment. In a nutshell, it’s the difference between mind processing and body reaction. A vision-language model with billions of parameters has to encode several camera images, run them through dozens of transformer layers with the instruction, and decode an action. On the compute a robot can actually carry in its chest, that takes something like 100 to 200 milliseconds. So it can produce maybe 5 to 10 decisions per second. But physics doesn't wait for that long. If a humanoid's foot slips or its ankle torque is slightly wrong, the fall develops over tens of milliseconds. To correct it you need to sense and respond every one or two milliseconds. That's 500 to 1000 times per second, or about 100X faster than the model takes to think. Same for grasping: the moment a finger touches an object, forces spike, and if you don't modulate within a few milliseconds you either crush it or drop it. You get an inversion of authority: the system component that understands the task can't react in time, and the component that can react in time doesn't understand the task.

- Limitation of simulation.

Sim-to-real works great for locomotion, because rigid-body contact with the ground is something physics engines model reasonably well. However, it works poorly for dexterous manipulation, because grasping involves soft contact, friction, deformable objects, and multi-finger force closure — all the places simulators are least accurate. That explains why you see so many robots walking, jumping, doing flips in the air, but you don’t see many robots using their hands to open a can of Coke. But it’s the hand where most of the economic value of a robot actually is.

- Compounding errors over long horizons.

Behavior cloning learns to imitate the demonstration distribution. The moment the robot’s actions and environment deviate slightly from that distribution, its observations become unfamiliar, its next action gets worse, and the deviations accelerate. Humanoid tasks are typically long; for instance, "unload the dishwasher" is dozens of sequential sub-behaviors, each of which has to end in a state the next one recognizes. Success rate compounds multiplicatively: eight steps at 95% each gives you 66% overall. That might not be sufficient for customers to pay for.

- Failure recovery might be missing from the data.

Human demonstrators are competent. They rarely drop things, and when they do, the take usually gets discarded. So the training set is full of successful execution and nearly empty of failure recovery, which means the robot has no idea what to do once something goes wrong, which is precisely when you need it to know. Deliberately collecting failure-and-recovery data is expensive and unnatural for teleoperators. Another issue arises when exploration of failure modes means breaking hardware or destroying the environment. Those instances often require a human to stand the robot back up and rearrange the scene. It’s very different from reinforcement learning deployed for software AI, and, you can imagine, very expensive.

- Knowledge transfer is imperfect.

Training a model on one robot doesn't automatically make it work on another. The robot's body is the interface to the world, and every body is different. A humanoid trained on one set of joints, motors, and sensors will have a different experience than a humanoid with a different set. The model has to learn how to map its understanding of the world to the specific kinematics and dynamics of its own body.

The field's main data hedge is pooling across robot types, but a humanoid's kinematics, workspace, and dynamics differ enough that knowledge transfer can only be partial. Even two humanoid models from different vendors have different joint limits, mass distributions, and hand designs. Some knowledge can transfer more easily (semantics, task structure), but the low-level motor mapping largely doesn't. Every hardware revision risks partially invalidating the data collected for the model training with the older generation hardware.

This is exactly the problem the robotics field is trying to solve. Rather than accepting that intelligence is trapped inside a specific body, some companies are trying to decouple the intelligence layer from the hardware entirely. For example, Skild AI is building a single foundation model on data pooled across many robot morphologies (quadrupeds, humanoids, arms, mobile manipulators, and human video) to create "omni-bodied" Skild Brain, an intelligence layer that can power different robotics applications. Pooling data across many bodies becomes a scalable solution for the robotics data problem. If that approach works, and Skild AI has published early demonstrations of this working across body types, the diversity of hardware stops being a tax on the data and becomes the thing that feeds the flywheel.

Why is this a strategic finance problem?

Every problem mentioned above is eventually requiring a decision about where to spend a finite amount of capital. And the humanoid industry has a structural bias toward spending it in the wrong place.

For instance, let’s talk about gravity again. Simulation is cheap for locomotion and nearly useless for manipulation. So the cheapest data produces walking, jumping, and backflips — the things that make a compelling demo and raise the next round. But the true economic value of a humanoid sits in the hand, in reliable manipulation of objects with valuable use cases, which is exactly the data that has to be manufactured slowly and expensively, commonly through teleoperation and real robot time. A company optimizing for what it can show on stage and a company optimizing for what a customer will pay for are pointed in different directions. Left alone, engineering and marketing pull toward the demo. The job of strategic finance is to resist the temptation and to influence capital allocation towards the boring but more commercially oriented data collection and training, because that is what eventually will push the company from being an entertainment company to a commercially viable business.

Willingness to pay for a physical task is a step function tied to reliability metrics: if a humanoid can accomplish the task only 66% of the time, customers are highly unlikely to pay for it. It has to perform with an acceptable level of reliability to be deployed in the real world and paid for. This means there is a wide valley between the impressive demo and the first payable task, and a company has to cross that valley on investor capital with no revenue underneath. What commercializable use cases to target first, what hardware setup that will require, how much capital it takes to achieve the reliability threshold, what’s the path to expanding the scope of capabilities in a capital-efficient manner should be amongst the most important questions on the C-suite agenda, and these are finance questions before they become engineering ones. Market sequencing is a capital allocation problem, the one that strategic finance must be tasked to evaluate before the engineering roadmap is locked in.

The unit economics also doesn’t behave like it does in pure software products, and pretending that it does is how these companies die. In software, the marginal cost trends toward zero and scale turns into margin. Here it doesn’t always behave like that: failure-and-recovery data has to be deliberately produced and never stops being expensive, new use cases might require new data collection efforts or new hardware support, every hardware revision partially invalidates the data behind it, and the sensor stack that generates the training data is itself a cost and a constraint. On the hardware side, a two-to-four-hour battery is a hard ceiling on utilization, and utilization is a key revenue driver for a robot. The number the buyer actually computes is total cost of ownership, including purchase, power, maintenance, warranty, replacement cycles, not the sticker price quoted for a demo robot. A $6,000 humanoid that falls off that stage is not a $6,000 product. The sticker price is a number that has to be reconciled against a bill of materials, a cost to serve, and maintenance cycles before it means anything. And that, again, is where strategic finance must be heavily involved to create a full economics profile and value proposition for the customer in order to guide pricing and engineering decisions.

Decoupling intelligence from hardware isn't only an engineering choice; it's a choice that impacts capital structure, unit economics profile, and value pool allocation. A company that builds the brain independent of the body can amortize combined data-collection cost across many robot platforms and customers instead of funding new model training with every hardware revision. It also can choose to push the costs of hardware, maintenance, and operations to their customers, supplying and monetizing the intelligence layer only. That is a more attractive capital story than vertical integration, of course assuming that the generalization across different hardware architectures actually holds. If it does work, there are still many strategic finance questions to address, including data economics, value share framework, pricing structure, integration / adaptation costs recovery, customer support framework, etc.

So what does strategic finance actually need to do? It needs to decide which task to commercialize first, mapping the reliability thresholds against customer willingness to pay. It needs to allocate capital across the data strategy — real robot time vs. teleoperation vs. simulation vs. human video, as each one has a different cost structure and risk profile. It needs to choose the business model — outright sale, robotics-as-a-service, outcome-based pricing — which is really a decision about who carries the utilization and reliability risk as the technology matures. It needs to establish the pricing guidelines that corresponds with the value generation. It needs to build the investor and board narrative that models the unit economics honestly, instead of importing software assumptions that don't hold. And it needs to enforce the discipline of tying capital deployment to commercial-ready milestones rather than demo milestones.

None of these are questions engineering can answer, and none of them can wait until the technology is finished, because the technical choices and the financial ones are the same choices. The hand design sets the bill of materials and the data cost. The data strategy sets the capital profile. The first market sets the runway. A humanoid company that treats finance as the team that books the numbers after the fact will make its most expensive decisions blind.

The robots will make breakfast eventually. But at the cost of many companies failing to account for economic realities of hardware-software products or failing to run engineering choices through the lens of capital deployment optimization. The companies still standing to sell humanoids will be the ones that were as rigorous about the economics as they were about the engineering. And economic rigor is what strategic finance brings to the table.

August 2026

© 2026 Nikolay Marcmin