AI Tech Observer

Infrastructure

What Actually Makes AI Expensive?

How can a training run cost only a few million dollars while data-centre investment reaches tens of billions and API prices keep falling? The answer lies in the interaction between cost boundaries, workloads and service objectives.

By AI 科技观察23 min read
先进封装、内存、网络与冷却部件组成的无人物算力成本工作台
Source: ·

One company says that training a model cost only a few million dollars. Another announces plans to invest tens of billions of dollars in data centres over the coming years. At the same time, developers see API prices continuing to fall. All of these figures may be true, but they cannot simply be added together, nor do they contradict one another. They describe different things: a single compute run, an ongoing supply capability and a commercial price offered to customers.

The question, then, is not "how expensive is AI?" but "for a particular AI capability, over what period and against which service objectives, who paid for which resources?" During training, chip time has to be converted into a model that converges. During inference, the same scarce resources have to be converted into output that meets requirements for quality, latency and availability. Buying accelerators is only one line item. Memory, packaging, networking, power and cooling, utilisation, depreciation, engineering staff and failed work all change the numerator; the model, traffic and acceptance criteria change the denominator.

This distinction directly affects procurement. Treating one final training run as the full cost of development leaves failed experiments, data and staff out of the budget. Treating an API list price as the provider's cost leads to false conclusions about why prices are falling. Comparing GPU prices alone ignores whether HBM, packaging, networking and facilities can turn peak compute performance into completed work. A useful cost assessment must fix the boundary first and only then discuss unit economics.

Contents

  1. 1.1. First decide which bill you are calculating
  2. 2.2. How HBM, packaging and networking enter the cost beyond GPUs
  3. 3.3. Once machines start depreciating, utilisation and facilities determine unit cost
  4. 4.4. Inference cost is not a fixed per-token figure
  5. 5.5. API prices, provider costs and customer total costs follow three different curves
  6. 6.Sources and citations

1. First decide which bill you are calculating

At least five kinds of figure are routinely described as "AI cost".

The first is a final training run: the compute used for pre-training, context extension or post-training under an established model design, dataset and training plan. The second is full development from research to release. Alongside the final run, this includes architecture exploration, ablation experiments, failed runs, data acquisition and cleaning, research and engineering staff, evaluation and safety work. The third is a stable inference service, where hardware depreciation, rental, networking, energy, operations, redundancy and underused capacity continue to incur costs. The fourth is the external price of cloud or API services. Only the fifth is the customer's total cost of ownership, which also includes integration, internal labour, human review, error handling, migration and business failure.

These five boundaries do not require a company to maintain five unrelated sets of accounts. Their purpose is to prevent the same expenditure from being omitted or counted twice. A provider's server purchase, for example, is a capital investment that subsequently enters the cost of service through depreciation. An API bill already includes the portion of the provider's costs passed on to the customer, so the customer should not add the provider's entire equipment investment to its own TCO. Conversely, the customer's data governance, evaluation and human review for an API integration are not included in the per-token price and must be recorded separately. The layer to which a cost belongs depends on who owns the asset, who bears the risk and which period is being compared.

In practice, every figure can carry four labels: cost object, time period, unit of measurement and responsible party. A training retrospective might be labelled "a particular final run, from start to finish, GPU hours, model developer". An annual service budget might instead be "production inference, one financial year, accepted requests, business unit and platform team". Figures are directly comparable only when all four labels match. When they do not, the analyst should first build a bridge showing which items need to be added, removed or amortised over the same period, rather than forcing two sets of accounts into the same terms with a single exchange rate or unit price.

DeepSeek-V3 is a useful example of why the boundary matters. Its technical report lists 2.664 million H800 GPU hours for pre-training, 119,000 hours for context extension and 5,000 hours for post-training, for a total of 2.788 million GPU hours. At an assumed US$2 per GPU hour, the authors convert this into US$5.576 million. Immediately afterwards, the report explicitly excludes earlier research and ablation experiments involving architecture, algorithms and data. [1]

The US$5.576 million figure therefore answers a question about the estimated cost of a set of final training activities. It is not the full bill for DeepSeek to develop this model capability. The dollar figure also depends on the authors' assumed rental rate; the GPU hours for each stage are closer to the underlying disclosure. It is reasonable to use the figure to compare the efficiency of final runs. Extending it into a complete cost for the research team, data, failed experiments and infrastructure crosses the boundary drawn by the report itself.

Epoch AI's study of a small number of frontier models likewise models the final training run, supporting experiments and R&D staff separately. It also distinguishes amortised hardware and energy, cloud rental and a fuller development cost. In the models it selected, the study assigns hardware, energy and R&D staff to separate cost categories. But the sample is limited and many inputs come from public disclosures and estimates, so its estimates do not constitute accounting ratios that can be applied directly to every company. [2]

Comparing two training-cost figures therefore requires a line-by-line check. Does each include only the final run? Does it include failed work? At which layer are staff and data counted? Is hardware measured by rental cost or depreciation? Are energy, financing and opportunity cost included? Without answers to those questions, two dollar figures precise to the last decimal place may still be incomparable.

The boundary also determines how failure is accounted for. A failed training run should not enter the denominator of "successfully trained tokens", but its cost cannot disappear from the numerator. An inference request that does not meet the quality threshold should not count as useful output, yet the compute, network and labour it consumed remain expenditure. The denominator most useful for business decisions is often neither GPU hours nor tokens, but "useful work that meets the defined quality and service objectives".

2. How HBM, packaging and networking enter the cost beyond GPUs

An accelerator's peak compute performance does not automatically become completed training or generation work. Memory has to hold model parameters, activations and KV state. Compute dies and memory have to be packaged into functioning devices. Once a task spans multiple accelerators, intra-system interconnects and inter-system networks determine whether data arrives in time. If any link leaves a processor waiting, an expensive compute unit may still be on the clock without producing a proportionate amount of useful output.

Official Blackwell architecture material shows two dies close to the reticle limit joined into a single GPU by a 10 TB/s chip-to-chip interconnect. The GB200 NVL72, meanwhile, organises 72 GPUs, 13.4 TB of HBM3E, 576 TB/s of aggregate memory bandwidth and 130 TB/s of NVLink communication into a liquid-cooled rack. [3][4] These specifications show why a GPU cannot be treated as a standalone product detached from memory, interconnects and cooling. They do not, however, prove that any component represents a fixed share of system cost, nor do they guarantee that production workloads will achieve the performance multipliers advertised by the vendor.

HBM first enters the provider's product-cost boundary. NVIDIA's 2026 Form 10-K discloses that the company buys memory from SK hynix, Micron and Samsung. Its descriptions of cost of revenue and inventory cost include purchased memory, other components, manufacturing support and related expenses. [5] Read alongside the HBM3E configuration of the GB200, this accounting disclosure supports a limited but important conclusion: HBM is a real input to a deployable system and enters product cost through procurement or manufacturing. The sources do not price HBM separately from the complete GPU, rack or cloud service, so they do not support a claim that "HBM usually represents a particular percentage of total cost".

HBM also has capacity economics of its own. Micron defines HBM as 3D-stacked DRAM connected by through-silicon vias. Its Cloud Memory Business Unit covers HBM, DDR, LPDDR and other products, and is managed in terms of revenue, cost of goods sold and operating income. The company has also disclosed construction of an advanced HBM packaging facility in Singapore and expansion of DRAM and HBM capability in Taiwan. [6] This establishes that HBM requires wafer production, stacking, packaging, testing and capacity-expansion resources. It does not allow HBM's unit cost to be separated from the business unit's aggregate figures, nor does it demonstrate a persistent industry-wide shortage.

Advanced packaging enters the accounts in another way. NVIDIA lists wafer fabrication, assembly, testing and packaging as parts of its manufacturing process, explicitly discloses its use of CoWoS, and includes assembly, testing, packaging and manufacturing-support expenses in cost of revenue and inventory cost. [5] TSMC describes CoWoS and SoIC as advanced-packaging capabilities that integrate chips of the same or different types to increase compute density, energy efficiency and integration while reducing latency. It also discloses continuing capital investment in the relevant capacity. [7]

Together, these disclosures define two boundaries. For a chip or system supplier, packaging is a product-manufacturing cost. From the perspective of supply capability, it is also back-end capacity requiring facilities, equipment, yield ramp-up and capital expenditure. But "requires investment" does not mean packaging is the second most expensive component in every AI system, nor does every dollar of capital expenditure rapidly produce an equivalent quantity of qualified products. The sources provide neither a standalone cost for each Blackwell device or CoWoS package nor a single global measure of the gap between supply and demand.

Networking crosses both capital and operating boundaries. Meta's 2025 Form 10-K includes depreciation of servers, network infrastructure and buildings, as well as energy and bandwidth expenses, in the cost of revenue for delivering its products. The filing also discloses capital purchases, leases and long-term contractual commitments for network infrastructure. [8] For a service provider, networking may therefore appear as equipment purchases and depreciation, bandwidth operating expenditure, lease payments or contractual commitments for capacity secured in advance. Meta's figures cover company-wide technical infrastructure and do not isolate AI networking, so they cannot be used to calculate a network share for each model or each million tokens.

Google's TPU v4 paper offers a more narrowly bounded example. In the 4,096-chip system described by the paper, optical circuit switches account for less than 5% of system cost and less than 3% of power, while helping to improve system scale, availability and utilisation. [9] This example shows that a networking component with a relatively small cost share can still change the effective use of expensive processors through topology. It does not show that other GPU clusters, Ethernet clusters or cloud services share the same figure of less than 5%.

A procurement sheet should therefore contain more than "number of GPUs multiplied by unit price". At a minimum, it should record whether the usable memory and effective bandwidth per card can accommodate the target model; whether packaging, testing, yield losses and supply commitments are included in the quote; which topologies the intra-system interconnect and inter-system network use; and who bears the cost of switching equipment, bandwidth, leases and capacity commitments. The final step is to test the target workload, rather than substituting peak specifications for useful output.

These items also need to be tested in linked scenarios, rather than choosing the cheapest option for each in isolation. Insufficient memory may force a model to be partitioned across more devices, increasing both communication and the number of devices that must be scheduled. Greater rack density may change network-port, power-supply and cooling requirements. Different delivery dates may mean that a system with a lower cost on paper produces business output months later. Buyers need not declare in advance which item is most expensive, but they should test how changes in capacity, bandwidth, delivery and utilisation alter the final denominator. The result is the system economics of the target workload, not a static ranking of component prices.

3. Once machines start depreciating, utilisation and facilities determine unit cost

The difficulty with hardware cost lies not only in the purchase price, but also in time. Fixed assets depreciate according to the calendar; business output is produced by requests and tasks. Gaps in training queues, communication waits, recovery from failures, low-traffic periods for inference services and unsuitable batching all allow depreciation, rental and facility costs to continue while reducing useful output over the same period.

The cost of a unit of useful work can be understood as the sum, over a given period, of depreciation or rental, cost of capital, energy, facilities, networking, operations and losses from failure, divided by output that meets the quality and service objectives. This is a comparison framework, not a requirement for every company to use the same accounting formula. Its value is that it forces buyers to examine numerator and denominator together. A lower equipment price may shrink the numerator; higher utilisation may expand the denominator. Conversely, stricter latency and availability objectives may require spare capacity and raise unit cost.

Utilisation cannot be reduced to a single fleet-wide average. A device that appears busy on a monitoring chart may be performing useful computation, or it may be waiting for memory, communication, data or another device. A cluster average can also hide a handful of heavily loaded devices alongside a large number of idle ones. Cost-oriented measurement should separate scheduled maintenance, failures, gaps in queues, communication waits, unaccepted runs and output that meets acceptance criteria, and should examine each during peaks and troughs. Only then can management tell whether the problem is a genuine lack of capacity, poor scheduling or spare capacity that the service objective requires but that cannot be sold to other workloads.

Depreciation assumptions alone can materially change costs recognised during a period. In 2023, Alphabet extended the estimated useful life of servers from four years to six, and that of certain network equipment from five years to six. This reduced depreciation expense that year by US$3.9 billion and increased net income by US$3.0 billion. [10] Alphabet's 2024 Form 10-K continued to include depreciation of technical infrastructure, network capacity, energy and equipment costs in its costs. It also stated that servers and network equipment are generally depreciated over six years, with obsolescence and planned utilisation taken into account. [11]

This does not establish six years as the correct useful life for the industry. Extending an asset's accounting life reduces current depreciation, but does not automatically improve the old equipment's performance or guarantee that it can continue operating with competitive energy efficiency and maintenance costs. A self-built system therefore needs three tests at once: how long the equipment is depreciated in the accounts, how long it remains technically serviceable and when it becomes economical to replace it with a more efficient system. These three lives need not be the same.

Networking crosses the capital and operating accounts in a similar way. Meta classifies depreciation of servers and network infrastructure, energy and bandwidth as cost of revenue, and records network-related leases and long-term commitments. [8] This means that "renting cloud capacity avoids capital expenditure" does not mean network capacity has no cost. Instead, contracts redistribute costs and risks: the provider assumes or transfers responsibility for equipment and capacity, while the customer pays through prices, minimum commitments, reservation arrangements or a flexibility premium.

Electricity is more than the chip's nameplate power. Google defines data-centre PUE as the ratio of total facility energy to the energy used by IT equipment; total facility energy also includes non-compute overhead such as cooling and power distribution. [12] Even with the same server power draw, cooling method, power conversion and facility design can therefore change total electricity consumption. PUE also has clear limits: it measures facility overhead, not model quality, chip utilisation or the cost of each successful request.

The US Department of Energy, citing research from Lawrence Berkeley National Laboratory, reported that US data centres consumed about 176 TWh of electricity in 2023 and gave a projected range of 325-580 TWh for 2028. AI was one of several demand drivers identified in the report. [13] These figures cover all data centres. The increase cannot be attributed entirely to AI, nor can the figures be directly converted into the electricity cost of a particular service. They demonstrate another layer of constraint: when grid connections, power equipment, cooling and construction lead times limit new capacity, facilities are no longer a background condition assumed to exist once servers have been purchased.

The correct comparison between self-built infrastructure and cloud rental is therefore not "total equipment price versus hourly rent". A self-built option requires sensitivity analysis of utilisation, depreciation period, residual value, maintenance, networking, energy, facilities and delivery time. A cloud option requires assessment of the rental premium, elasticity, reservation commitments, speed of deployment and exit costs. Cost transfers become visible only when the same workload and service objectives are applied to both options.

4. Inference cost is not a fixed per-token figure

The end of training is not the end of the cost question. Inference is a workload that varies with the model, the request and the service objectives. The same one million tokens can place very different demands on compute, memory and scheduling depending on whether they come from short inputs with long outputs, long inputs with short outputs, low-concurrency real-time conversations or high-concurrency offline generation.

Model size first determines how many weights must be moved and how much computation must be performed. Sparse models also require a distinction between total parameters and the parameters actually activated for one request. Input length primarily affects the prefill stage, in which the system processes the existing context in parallel. Output is generated token by token through autoregressive decoding, with another round of computation required for each new token. The ORCA paper explains this iterative process and shows why fixed whole-request batches make completed requests wait while new requests queue. [14]

Batching allows multiple requests to share weight reads and device execution, increasing throughput, but it conflicts with queueing time and tail latency. Offline tasks with sufficiently steady traffic and tolerance for waiting can form larger batches. Low-traffic or latency-sensitive real-time services may be unable to fill the same batch for long periods. A "per-token cost" should therefore be accompanied, at a minimum, by concurrency, batch size, time to first token, time per output token and tail latency. Otherwise, a low price may simply have been bought with a longer wait.

The KV cache is a second, often underestimated constraint. It stores attention keys and values already computed during generation, grows with concurrency and sequence length, and occupies accelerator memory. The vLLM paper notes that each request's KV cache is large and changes dynamically, while memory fragmentation and duplication restrict attainable batch size. In the paper's tests, its PagedAttention method achieved two to four times the throughput of the selected baselines at the same latency. [15] This multiplier applies only to the particular models, hardware, sequences and decoding settings tested; it is not a cost reduction that every production system can realise. The result shows that memory management changes utilisation and throughput; it does not establish a universal savings rate.

Prefill and decoding may also compete for different resources. DistServe separates time to first token (TTFT) from time per output token (TPOT), describes the interference that arises when both stages are colocated, and measures service capacity by the request rate that meets both latency objectives. [16] Google's inference research likewise treats partitioning, batch size, model FLOPS utilisation, memory requirements, context length and strict latency as parts of one optimisation problem. [17] This explains why a configuration that performs efficiently in an offline throughput test may not suit an interactive product requiring a fast first token and stable tail latency.

Caching also involves two distinct mechanisms. A request's KV cache stores the state of that generation; cross-request prompt caching attempts to reuse an identical prefix. Anthropic's documentation explains that prompt caching depends on an identical prefix and a cache hit. On a miss, the full prompt must still be processed before a reusable entry is created. [18] Fixed system prompts, large repeated documents or templated calls may therefore benefit. Traffic whose prefixes change frequently cannot treat an advertised caching discount as a dependable saving.

Reliability further changes capacity requirements, but the available sources cannot provide a redundancy ratio or retry cost that applies to every service. Replicas, multi-region deployment, capacity reservation, rate limiting, fallback and retries should be treated as a procurement checklist, not as a uniform cost rule established by the research literature. During a pilot, a company needs to record each item: the target availability and tail latency; how the system applies rate limits under overload; whether it retains replicas or multi-region capacity; which model it falls back to; when it triggers retries; how failures and retries are billed; and whether the result after a retry still meets the acceptance criteria.

The final inference test sheet should include, at a minimum, the distributions of input and output tokens, concurrency, batch size, KV-cache occupancy, cross-request cache hits, TTFT, TPOT, tail latency, success rate, retry rate and quality. Its cost denominator should be "accepted results", so that a system generating large volumes of useless output at a lower unit price is not mistaken for an efficiency improvement.

Testing should also cover traffic patterns rather than running a benchmark once at a stable full load. A company can construct separate scenarios for normal load, sudden peaks, long contexts, long outputs, high cache-hit rates and low cache-hit rates, then observe how throughput, latency, failures and the bill change at the same quality threshold. If an option shows a low unit cost only under sustained full load, the next question is whether the actual workload can sustain batches of that size. If a low-latency option depends on substantial idle headroom, that capacity should be explicitly assigned to the cost of the service objective. Scenario results are more useful than a single average when deciding model routing, capacity reservations and the boundary between owned and rented infrastructure.

5. API prices, provider costs and customer total costs follow three different curves

The API list price is the price a customer will actually pay, but it is not the provider's cost statement. It is usually charged by input, output, caching or another service unit. Behind it may sit hardware and energy, engineering and operations, capacity risk and commercial decisions, but a public price list alone cannot identify the contribution of each factor or reveal the provider's profit margin.

Epoch AI's study of inference-price trends tracks the lowest API price that reaches specified benchmark thresholds and uses a weighted input-output price. The study states explicitly that it did not model cost drivers and lacked sufficient evidence to determine whether price declines reflected narrower profit margins. [19] The article can therefore say that "an observed decline in API prices is not the same as an identified decline in provider costs", but it cannot use this evidence to construct a definitive causal table for prices.

Customers need another set of accounts. A practical total-cost-of-ownership framework has four layers. Direct expenditure includes API or cloud bills, storage, retrieval, tool calls, networking, logging and capacity purchased to meet service objectives. Internal labour includes integration, maintenance of data and prompts, evaluation, security and compliance, human review and production operations. Losses from failure include failed runs, retries, rework, incidents and measurable losses caused by missed service objectives. Opportunity cost includes waiting, migration, engineering capacity tied up and other opportunities forgone because of delayed delivery.

These four layers are an accounting tool proposed by this article, not an industry standard established by the sources. Nor can they be added together unconditionally. They are suitable for aggregation into the TCO of a business result only when the time period, quality threshold, measurement basis and cost ownership are aligned. Opportunity cost, in particular, should be reported separately from cash expenditure. Otherwise, potential revenue loss in one estimate may be mixed with a cash bill in another, producing a total that is precise but impossible to interpret.

Efficiency improvements still matter. The Stanford 2025 AI Index reports that the inference price of a system performing at GPT-3.5 level fell by more than 280-fold between November 2022 and October 2024. [20] This comparison is based on a particular performance threshold. It does not mean every task received equivalent value, nor does it by itself establish the cause of the fall in underlying cost or price. It does show that the price of a unit of capability can change sharply over a relatively short period.

A lower unit price does not guarantee lower total expenditure. Cheaper calls make previously uneconomic features viable. Products may use longer contexts, generate more candidates, add agent steps or embed models in more processes. If call volume and the amount of work per call grow faster than unit cost falls, total demand and total expenditure will still rise. Conversely, if task volume is fixed, output is controlled and efficiency gains do not induce more use, total expenditure may indeed fall. Price trends and trends in total data-centre electricity use can jointly demonstrate this possibility, but they cannot prove a causal relationship between the two. [13]

Procurement decisions can be reduced to six questions:

1. What is the boundary: a final training run, full development, a provider's product and capacity, a stable service, or customer TCO?

2. What is the denominator: GPU hours, tokens, requests, or accepted business results?

3. What is the workload: how large is the model; how long are inputs and outputs; and what are the concurrency, batch size, cache-hit rate and peak-to-trough distribution?

4. What are the service objectives: quality, TTFT, TPOT, tail latency, availability and data boundaries?

5. Who bears fixed costs and responsibility for reliability: do memory, packaging, networking, depreciation, idle capacity, reservations, redundancy and retries fall to the provider, the self-hosting customer or both parties through the contract?

6. Which items are not yet in the quote: have data, integration, retrieval, evaluation, human review, failure, migration and exit costs been recorded in separate layers?

This method will not produce an "average price of AI" detached from context. It makes the real changes visible. At times, a chip or API unit price falls while costs move into networking, service objectives and internal engineering. At other times, better scheduling, caching and model choice increase useful output and genuinely improve unit economics. Only after fixing the boundary, denominator and acceptance criteria can a buyer judge whether a cost is falling, being transferred or being amplified by greater demand.

Sources and citations

  1. DeepSeek-V3 Technical Report

    DeepSeek-AI / arXiv · Table 1; Introduction

  2. How much does it cost to train frontier AI models?

    Ben Cottier, Robi Rahman, Loredana Fattorini, Nestor Maslej and David Owen, Epoch AI · Executive summary; Figures 1-3; Methodology

  3. NVIDIA Blackwell Architecture

    NVIDIA · NVIDIA · Blackwell Architecture; Fifth-Generation NVLink

  4. NVIDIA GB200 NVL72

    NVIDIA · NVIDIA · Specifications

  5. NVIDIA Corporation 2026 Annual Report on Form 10-K

    NVIDIA / U.S. SEC · Item 1, Manufacturing; Item 7, Gross Profit and Gross Margin; Note 1, Inventories; Item 1, Manufacturing; Item 7, Cost of revenue; Note 1, Inventories

  6. Micron Technology, Inc. 2025 Annual Report on Form 10-K

    Micron / U.S. SEC · Item 1, Products and business units; Item 7, Manufacturing investments; Segment disclosures

  7. Taiwan Semiconductor Manufacturing Company Limited 2025 Annual Report on Form 20-F

    TSMC / U.S. SEC · Item 4, Information on the Company; Item 5, Operating and Financial Review

  8. Meta Platforms, Inc. 2025 Annual Report on Form 10-K

    Meta / U.S. SEC · Item 7, Cost of revenue and investing activities; Notes 1, 6, 7 and 11; Item 7; Notes 1, 7 and 11

  9. Alphabet Inc. 2023 Annual Report on Form 10-K

    Alphabet / U.S. SEC · Note 1, Change in accounting estimate

  10. Alphabet Inc. 2024 Annual Report on Form 10-K

    Alphabet / U.S. SEC · Cost of revenues; Note 1, Property and equipment

  11. Data center efficiency

    Google Data Centers · PUE methodology

  12. DOE Releases New Report Evaluating Increase in Electricity Demand from Data Centers

    U.S. Department of Energy · U.S. Department of Energy · DOE release summarising LBNL report

  13. Orca: A Distributed Serving System for Transformer-Based Generative Models

    Gyeong-In Yu et al., USENIX OSDI 2022 · Abstract; Sections 2-3

  14. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

    Yinmin Zhong et al., USENIX OSDI 2024 / arXiv · Abstract; Sections 1-2

  15. Efficiently Scaling Transformer Inference

    Reiner Pope et al., arXiv · Abstract

  16. Prompt caching

    Anthropic Claude Platform Docs · How prompt caching works

  17. LLM inference prices have fallen rapidly but unequally across tasks

    Ben Cottier, Ben Snodin, David Owen and Tom Adamczewski, Epoch AI · Methodology; Key findings and limitations

  18. The 2025 AI Index Report

    Stanford Institute for Human-Centered Artificial Intelligence · Stanford HAI · Top Takeaway 7