OpenAI wants to control the full stack, from data centers to ChatGPT

OpenAI connects chips, data centers, models, and products to reduce the cost of each task completed by ChatGPT, Codex, and its agents.

OpenAI no longer wants its models to depend on any single layer of the technology industry. In an article by CFO Sarah Friar, the company lays out a strategy spanning data centers, chips, networks, models, its developer platform, and consumer products.

The plan is to optimize these components together. A more efficient model reduces the number of attempts required to finish a task. Better serving software cuts unnecessary computation. A chip designed for OpenAI’s workloads delivers faster responses. Those gains should then appear in ChatGPT, Codex, the API, and OpenAI’s future devices.

Jalapeño, its first custom inference chip, gives this strategy a physical form. Announced in June and developed with Broadcom and Celestica, it was designed around the requirements of large language models rather than adapted from a general-purpose accelerator. OpenAI controls the architecture, while its partners handle silicon implementation, networking, boards, and rack integration.

The first measured results cover three open models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. According to OpenAI, Jalapeño delivers 1.5 to 1.9 times more work per watt at peak throughput than the comparison systems, with end-to-end latency reduced by a factor of 1.7 to 3.6. For highly interactive workloads, the claimed performance advantage ranges from 2.1 to 4.1 times.

The tests use InferenceX, a public benchmark developed by SemiAnalysis. GPT-OSS 120B is compared with an NVIDIA GB200 system, while DeepSeek R1 and Kimi K2.5 are tested against the GB300. Jalapeño has a rated power envelope of 700 watts, compared with 1,200 watts for the GB200 and 1,400 watts for the GB300. OpenAI says its sustained consumption remained at or below 550 watts during testing.

These figures do not yet provide a complete comparison of operating costs. OpenAI normalizes the results using each accelerator’s published power rating without including total data center consumption, purchase prices, hosting agreements, software maturity, or availability at scale. The company published the results using a public protocol, but they remain specific to the models, number formats, and configurations selected.

The custom chip is not intended to replace OpenAI’s entire existing fleet. The company continues to describe Microsoft and NVIDIA as the two historical foundations of its infrastructure. Its supplier portfolio now also includes AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, and SoftBank. Each contributes to a different part of the system, from cloud infrastructure and accelerators to energy and data center development.

The goal is to preserve several options. The most expensive systems can be reserved for frontier-model training or workloads that require maximum capability. Large and predictable inference jobs can be directed toward hardware with more favorable economics. Jalapeño adds an internal option to this portfolio and gives OpenAI greater bargaining power with suppliers.

This integration therefore remains selective. OpenAI builds components when designing them alongside its models could produce an advantage, but it continues to buy elsewhere when partners have stronger technology or can deploy it faster. The “full stack” described by Sarah Friar is not complete self-sufficiency. It is the ability to distribute workloads and investments across multiple systems.

Data centers sit at the other end of the strategy. Project Camellia, planned for Effingham County, Georgia, is expected to receive 3.2 gigawatts of power in phases between 2028 and 2032. OpenAI says it will cover the necessary infrastructure and energy costs, use a closed-loop water system, and subject its commitments to an annual public audit by an independent firm.

Cost control does not depend on hardware alone. OpenAI also cites GPT-5.6 Sol, which scored 80 on the Artificial Analysis Coding Agent Index. Artificial Analysis’s evaluation confirms that the model, running in Codex at maximum reasoning effort, leads the three coding tests included in the index. OpenAI places particular emphasis on how much text the model generates to reach that result, which is lower than several similarly capable models.

That reasoning changes the economic unit being measured. The price of one million tokens says little about the cost of an agent if it repeats attempts, produces unnecessarily long answers, or fails before completing the task. OpenAI therefore prefers cost per successful task, a measure that includes retries, execution time, and the resources consumed across the entire workflow.

Greater efficiency does not guarantee lower total consumption. Friar invokes the Jevons paradox herself: when a technology becomes cheaper, new uses become economically viable and demand increases. Companies can analyze more documents, run more simulations, or assign longer processes to agents. Each task costs less, but the number of tasks grows.

The article ultimately presents a financial strategy as much as a technical one. OpenAI wants to lower its serving costs, avoid excessive dependence on any single supplier, and convert those savings into additional capacity. Jalapeño is the first internally designed component in that structure, not its final form. Its value will now depend on large-scale production, reliability, and the complete cost of the systems in which it operates.