Cerebras Systems Q2 2026 Earnings Call Transcript

Key Takeaways

  • Positive Sentiment: Q2 exceeded guidance, with core revenue of $209.9 million, up 103% year over year, core gross margin of 40.6%, and core operating margin of negative 16%.
  • Positive Sentiment: Management raised its full-year 2026 outlook to core revenue of $880 million–$890 million, gross margin of 41%–43%, and operating margin of negative 19% to negative 17%; it expects Q3 gross margin to be the low point before improving in Q4.
  • Positive Sentiment: Cerebras reported $25.4 billion of remaining performance obligations and expects to more than triple core revenue in 2027, supported by more than 600 megawatts of data-center capacity secured or under contract and manufacturing capacity expected to increase more than tenfold in 2026.
  • Positive Sentiment: The company is expanding its inference opportunity through disaggregated solutions with AMD and AWS, combining GPU or Trainium prefill with Cerebras decode; management said the AMD configuration can preserve Cerebras speed while increasing throughput by up to 5x.
  • Neutral Sentiment: AWS deployment is expected to become generally available through Amazon Bedrock in Q1 2027, while additional hyperscaler revenue is expected from mid-2027; management acknowledged that OpenAI will remain a meaningful customer but should decline as a percentage of revenue over time.
AI Generated. May Contain Errors.
Earnings Conference Call
Cerebras Systems Q2 2026
00:00 / 00:00

There are 10 speakers on the call.

Operator

Good afternoon and welcome to the Cerebras Systems second quarter fiscal year 2026 earnings conference call. At this time, all participants are in a listen-only mode. Following management's prepared remarks, we will open the call for questions. Please note that today's call is being recorded. I will now turn the call over to Sean Dorsey, Head of Investor Relations. Please go ahead.

Speaker 1

Thank you, operator. Good afternoon, everyone, and welcome to Cerebras Systems Q2 2026 earnings call. Earlier today, we issued our press release and posted our supplemental earnings presentation to the investor relations section of our website. A replay of this webcast will also be available on our investor relations website following the call. Joining me today are Andrew Feldman, our Co-founder, Chief Executive Officer, and President, and Bob Komin, our Chief Financial Officer. Before we begin, I would like to remind everyone that today's discussion will include forward-looking statements under the safe harbor of the Private Securities Litigation Reform Act of 1995. These statements include, but are not limited to, statements regarding our future financial performance, business strategy, market opportunity, customer demand, product roadmap, technology leadership, supply chain, operating model, and outlook for Q3 and full year 2026.

Speaker 1

Forward-looking statements are based on our current expectations and assumptions and are subject to risks and uncertainties that could cause actual results to differ materially from those expressed or implied. These risks are described in our SEC filings, including our final prospectus related to our IPO and our future periodic filings with the SEC. We undertake no obligation to update these forward-looking statements except as required by law. During today's call, we will also discuss certain non-GAAP financial measures. Reconciliations between GAAP and non-GAAP results are included in today's press release and supplemental materials, which are available on the investor relations page of our website. With that, I'll turn the call over to Andrew.

Speaker 2

Thank you, Sean. Thank you all for joining us today. Q2 was a strong quarter. We completed our public offering, but we did not let that distract us from execution. We delivered record core revenue and beat guidance on all metrics: core revenue, core gross margins, and core operating margin. Looking forward, we see unbound demand for fast inference. The market is realizing that speed is not a benchmark item. Speed changes user engagement, it changes agentic performance, and it changes AI productivity. Fast inference unlocks new applications and new markets. As we've shared with you previously, 2026 is a foundation-building year for Cerebras. We've made excellent progress on multiple fronts in the past seven weeks since our last earnings call, preparing us for a massive 2027, 2028, and 2029 as we deliver on the $25 billion of RPO we currently have on our books.

Speaker 2

With the benefit of that progress, we expect to more than triple our core revenues in 2027 and continue to grow at multiples in the years following. We think of progress in terms of capacity, capabilities, and customers. We are expanding capacity by adding new contracts for data centers around the world, expanding manufacturing capabilities, and collaborating with our vendors to ensure supply and to support our extraordinary growth. We are advancing our capabilities by inventing new technology that extends our performance and throughput and our power efficiency. We are expanding our customer base by accelerating AI productivity in existing markets like coding and agentic flows, and pioneering new areas like security, where speed opens up entirely new opportunities. On the capacity front, data center space continues to be the bottleneck for the entire industry, and we are no exception.

Speaker 2

The faster we and our customers bring on new data centers, the faster we grow. Over the last seven months, we have had an all-out push to secure and build out data centers. We have two advantages. First, because we are serving inference, we do not need gigawatt footprint locations like those needed for training clusters. This gives us much more flexibility to scale up capacity across a multitude of locations around the world. Second, we have built a repeatable process for site selection, cluster deployment, and customer activation, which is an operational muscle required to turn gigawatts into production tokens at a global scale. I am pleased to report that our push has been very successful. We now have data centers either up or under contract in Alabama, Dallas, Denver, Minneapolis, Santa Clara, Stockton, and outside the U.S. in France, Finland, Manitoba, Montreal, Norway, Saskatchewan, and Toronto.

Speaker 2

In total, over the last seven months, we have secured more than 600 megawatts of data center capacity that is either live now or will be delivered by the end of 2027. While this is not nearly enough to meet our demand, our data center pipeline of new opportunities for expansion continues to grow and is now measured in gigawatts. To put it in perspective, as we continue to build out our first-party cloud, it will be among the largest non-hyperscale AI clouds. Whereas at the end of 2025, we were on a steep learning curve, today I am happy to report that we are pretty good at data center build-out with a clear path to becoming excellent. Other key dimensions of capacity include manufacturing and supply chain.

Speaker 2

Here we have successfully increased our manufacturing capacity and are building up new factories with Flex in Sanmina, and expect to increase our manufacturing capacity by more than 10x in 2026 and continue that expansion in 2027, again, preparing us for the exceptional growth expected in the years ahead. Our partnership with our supply chain vendors has also turned into a significant advantage. TSMC has once again come through, and we have the wafers needed to fuel our growth. Our ability to get wafer supply also benefits from the fact that we were able to deliver industry-leading performance while running on TSMC's five nanometer node, where wafers are less expensive and supply is less constrained. Our decades-long relationships with our supply partners reinforces our confidence that we can deliver on our growth plans going forward. These relationships are rare and valuable, particularly in times of short supply.

Speaker 2

Finally, recall that most of the critical supply chain constraints currently faced by the industry don't apply to us. For example, we don't use HBM memory, CoWoS packaging, or require 3 nanometer fab capacity. On the capabilities front, in the second quarter, we delivered support for OpenAI's GPT-5.6 Sol, the largest and most capable of the frontier models. In fact, Cerebras serves 5.6 Sol at a speed that is 10X faster. With GPT-5.6 Sol, this lays to rest any of the remaining concerns regarding our ability to support large frontier models. Being a partner for the delivery of GPT-5.6 Sol and serving it to our cloud speaks to the maturity of our software stack. It takes millions of system hours of production hardening to get to the point where one can deliver hyperscale quality and reliability.

Speaker 2

We're proud that our inference cloud can meet the requirements of the most demanding customers. Our collaboration on serving models at the frontier has opened up new and significant strategic advantage previously only available to NVIDIA. Closed-source frontier models include a continual stream of new insights and new AI techniques. Serving these models allows us to see into the future and to prepare for it. Our roadmap from the hardware through the software stack now reflects what we're seeing and will give us a compounding advantage in the years to come. Continuing on the theme of capabilities, let's turn to disaggregation. We now have disaggregated inference solutions with two of the leading chip companies, AMD with their Helios and AWS with Trainium. Disaggregation expands the market for both the GPU provider and for Cerebras. Disaggregation enables GPUs to participate in a market currently foreclosed to them, namely fast inference.

Speaker 2

Disaggregation enables Cerebras to expand our opportunity to those customers who are more price sensitive and expands the profitability of our data centers. Let's see how this works. As with any compute market, as inference grows and matures, opportunities for specialization emerge. Disaggregation is a form of specialization that is particularly well-suited for workloads with well-known traffic patterns. In these cases, disaggregation delivers advantage by separating inference into two stages, prefill and decode, and using different processors for each stage. Prefill processes the input from the user or agent. It is a parallelizable workload. As a result, prefill is well-suited for GPUs and their HBM-based memory architectures. Decode generates the output tokens. It's the harder technical problem and is the bulk of the computational work in a disaggregated solution. It is sequential and memory bandwidth intensive, and is particularly well-suited for our Wafer-Scale Engine.

Speaker 2

The prefill and decode processors need to be linked to create the end-to-end solution, and this is where standards-based IO and open engagement strategy has made integration easy and straightforward for Cerebras. A few weeks ago, we announced a partnership with AMD to build disaggregated inference solutions. The solutions combine their Helios racks with our CS systems. The combined solution maintains Cerebras' speed while increasing throughput by 5X. To understand how powerful this is, it's important to understand the difference between speed and throughput. Speed is a measure per user. It's measured in tokens per second per user. It is how fast your query is answered, or how long it takes an agent to finish a task. Here, it is on the x-axis. Throughput, on the other hand, is the total number of tokens the solution can produce per second.

Speaker 2

It is measured by adding up all the tokens across all the simultaneous users. Here, it is shown as it is generally done on the y-axis. Speed is critical for user experience. Throughput is critical for inference economics. GPU solutions can support high throughput, but only at low speeds. When configured to support even moderate speeds, GPU throughput drops precipitously. This is true not just for GPUs, but also for ASICs and all solutions that use HBM. The HBM memory architecture forces a trade-off between throughput and speed. SRAM-based architectures like Cerebras' are the exact opposite. We support blisteringly fast tokens, but at moderate throughput. GPUs want to get faster without giving up throughput. Cerebras wants more throughput without giving up speed. Herein is the strength of our disaggregated solution. It delivers Cerebras speed with 5x higher throughput.

Speaker 2

Increasing throughput by 5x while keeping our industry-leading speed has a profound impact on the economics of token generation. It means up to 5 times as many high-speed, high-value tokens are made by each Cerebras system. More tokens per system at lower cost means more revenue and more gross margin. More tokens generated per CS system also means more tokens per watt, making each data center more profitable. Perhaps most important in a data center-constrained environment, the disaggregated solution allows us to serve more of the demand that we have in RPO. Finally, we believe this disaggregation approach makes performance and economic sense with any GPU.

Speaker 2

For operators who have already deployed large footprints of GPUs, disaggregation with Cerebras offers them an opportunity to create meaningful leverage built on their existing investments by pairing some portion of those GPUs with Cerebras solutions, dramatically improving the value and usefulness of their data center footprint. Continuing on the capabilities theme, let us turn to our roadmap. Our engineering execution is continuing at pace. We expect to deliver new systems that double our speed each year for the next several years. Remember, we are doubling our performance, starting with a 15x performance advantage over everyone else in the industry. In addition, while keeping the performance crown, over the next 18 months, we plan to deliver solutions that increase throughput by more than 20x. Next week at our annual Supernova conference, we will be unveiling the CS-4, our fourth generation system.

Speaker 2

It will be a great event with lots of product announcements, so I recommend you attend. Finally, we are currently on track to launch our CS-5 in the second half of 2027. Looking even further out, our invention engine is humming. We have significant partnerships with the U.S. government for delivery of stacked memory solutions, as well as integrated wafer scale optical solutions. In the years ahead, you can expect to see inventions from us in chip and chip architecture, as well as all elements of system design, including packaging, IO, and power delivery. To summarize the capabilities section, we expect to continue to deliver pioneering advances in product and technology to drive up speed and throughput, reduce the power used per token, and slash the cost per token of our solution. Now let us turn to the customer front.

Speaker 2

Fast tokens are in demand and command a premium at market, and fast tokens with frontier intelligence are only available through OpenAI Cerebras partnership. Our work with AWS continues, and we expect to have solutions generally available in Q1 2027 through AWS's Amazon Bedrock platform. This AWS partnership expands our market opportunity and provides us with global reach through an industry leader who is trusted by nearly every enterprise in the world. Our discussions with other hyperscalers are also going well. We expect to produce first revenues starting in mid-2027 and ramp through 2028 and beyond. With all of this progress, I think it's important to keep in mind that our $25 billion in RPO does not reflect any backlog of business from AWS or any other hyperscaler at this time. Our business outside of OpenAI and the hyperscalers continues to grow nicely.

Speaker 2

For example, in Q2, we signed six deals north of $30 million. AI coding continues its rapid rate of growth. In our experience, no one says, "I'm happy with slow tokens" when coding. So not surprisingly, in the coding category, our footprint continues to grow. We signed new agreements with public companies such as Figma and startup leaders such as Cognition, and we extended our presence in Europe with a major win at Lovable. Agentic flows are growing quickly, and the value of speed compounds as agentic operations rapidly evolve toward multi-step, multi-agent solutions. Companies as diverse as Block, AlphaSense, and GSK signed new agreements during the second quarter with Cerebras to leverage fast inference to provide their custom-made agents. Fast AI also opens up new markets, extending the TAM for Cerebras. Security is one such example.

Speaker 2

Our recent win with CrowdStrike is an application that only exists if AI is fast. Fast AI enables AI-based security devices to sit in line with enterprise traffic and use LLMs to secure traffic so quickly that nobody notices. Fast AI enables an LLM to provide security that is invisible to users. The AI provides the security. The speed creates the invisibility that enables the security to avoid delay and disruption. We expect this type of security to become the norm given the rapidly evolving threat landscape. Enterprises will soon expect vast swaths of their traffic to be inspected in this way, creating massive new opportunities made possible exclusively through Fast AI. Frontier labs, hyperscalers, leading chip makers, the fastest-growing startups, and massive enterprises are all now customers and partners of Cerebras and benefit from our blazing fast inference. To summarize, overall a strong quarter.

Speaker 2

We went public in a successful IPO. We beat on all metrics, core revenue, core margins, and core operating margins. We made progress in each of our key domains, capacity, capability, and customers. These are the foundations on which we will achieve our goals of massive growth in 2027 and 2028 and continue this exceptional rate of growth in 2029 and beyond. With that, I'll turn things over to Bob. Bob?

Speaker 3

Thank you, Andrew, and good afternoon, everyone. We made tremendous progress in the first half of 2026. As we described, 2026 is the foundation for multiples of growth over the years ahead. We entered the year having won one of the largest technology deals ever, creating RPO of more than $25 billion. This required us to immediately work on major increases in three critical components of capacity. First, we needed to increase our wafer supply. As Andrew described, due to our strong relationship and support from TSMC, we did that and are now well positioned, not just for the remainder of this year, but for the next year as well. Second, we needed to scale our manufacturing capacity. We are already 4 times above where we were in the first half of 2025, and we will have increased the manufacturing capacity more than 10x in 2026.

Speaker 3

We are making great progress here. Third, we need to substantially increase our data center capacity. We have made significant progress with over 600 megawatts now up or under contract, expected for delivery by the end of 2027, plus a pipeline in gigawatts. The foundation is in place to support a tripling or better in our core revenue in 2027 and additional multiples in future years. This growth is also setting us up for significant margin expansion in 2027 and beyond. Turning to Q2 financial results. We had another quarter of strong results, beating expectations across each element of our guidance. We delivered record core revenue. We beat on core gross margin and on core operating margin. I will be using the same core business framework introduced last quarter to describe our progress.

Speaker 3

The definition of our core business metrics and reconciliations of all of them to GAAP are included in today's earnings release and on our website. Core revenue was $209.9 million, up 103% year-over-year. Our private cloud business is growing at an extraordinary pace. Core cloud and other services revenue was $127.7 million, up 287% year-over-year. This nearly fourfold increase reflects the tremendous demand we have for Cerebras' Fast Inference service. Core hardware revenue was $82.1 million in the quarter, up 17% compared to last year. We focus on total core revenue, not the mix between the two, which can vary significantly quarter to quarter due to the timing of large new cloud capacity additions and hardware shipments.

Speaker 3

In Q2, most of the total core revenue was attributable to increases in our core cloud offering, reflecting the ramp in our OpenAI deployment, increases in our other cloud customers' usage, and finally, hardware customers who are also wrestling with the timing of new data center capacity. The demand for Fast Inference continues to be strong, with several late-stage hardware deals representing hundreds of millions of dollars in the pipeline from new customers, as well as significant new cloud deals for 2027. Existing Fast Inference markets are growing, and new ones are getting started. We see disaggregation as an important unlock to drive new use cases since it dramatically improves the economics of inference and of data center ownership.

Speaker 3

Today, this means that up to 5X more tokens are produced per CS system, so power and cost are significantly reduced per token. By continuing to invest heavily in R&D and our product roadmap, Cerebras will quadruple our current industry-leading speed and increase throughput by more than 20X through the end of 2027, drastically improving our performance and the economics of inference. Turning now to gross margin. Year-over-year, core gross margins improved substantially. Q2 core gross margin was 40.6%, approximately 940 basis points higher than Q2 2025. This reflects the increase in value of fast inference by the market, our continuous stream of product improvements, and additional economies of scale. Breaking the total core gross margin into its components, core cloud and other services gross margin was 41.8%, 1,600 basis points better than Q2 2025.

Speaker 3

Core hardware gross margin was 38.8%, 510 basis points higher than a year ago. As we described last quarter, we are meeting some of the overwhelming demand for our fast inference service by temporarily renting some of our own systems back from our cloud customers and making it available through the Cerebras cloud. Serving this inference demand sooner strengthens our ability to meet the needs of our cloud customers and to grow with them over time. We believe this will create additional long-term value for Cerebras and its shareholders. In the short term, it reduces gross margin as we have a higher cost for this rented capacity. As a result, sequentially, core gross margin was 40.6% versus 46.5% in Q1 2026.

Speaker 3

Had we not had higher costs due to increasing our private cloud capacity by renting back more of our systems, core gross margins would have been approximately 500 basis points higher and more similar to last quarter. Looking forward, we expect Q3 to be the low point for core gross margin before improving significantly in Q4 2026 as we bring on more data centers filled with lower-cost Cerebras owned systems. This will cause core cloud gross margin to step back up. Core gross margin will also continue to improve in 2027 and trend towards our target of 60% plus for several reasons. The market has recognized that fast tokens are more valuable tokens. This supports higher pricing, which is reflected in hardware and cloud deals that will be recognized over the next several quarters.

Speaker 3

Over the next few quarters, we will roll off higher cost rented systems and replace them with lower cost owned systems in our private cloud. Our product roadmap has us increasing throughput by 20X over the next 18 months. This reduces the cost to produce tokens per system and per unit of power. As our scale grows, our bill of material costs and our supply chain will improve more. By being on the 5 nanometer node, our wafer costs are lower than others who need to be on the 3 or 2 nanometer node. We are also purchasing wafers in much higher volumes. Finally, we do not rely on HBM, which pressures those who use it to either raise prices or lose margin points. We are not exposed to that risk, which we believe will improve our value proposition and pricing flexibility. Turning to operating margin.

Speaker 3

Core operating loss was $33.6 million. Core operating margin was negative 16% compared to negative 42% a year ago, an improvement of approximately 2,600 basis points year-over-year. Our ability to deliver this significant improvement in core operating margin while more than doubling revenues and stepping up our investments in all areas demonstrates the strong operating leverage inherent in our business model. Today, we are investing in world-class people, manufacturing and data center capacity, and company infrastructure to support the significant increase in scale we expect to deliver over the next several years. Remaining performance obligations at June 30, 2026, are $25.4 billion. This backlog provides us visibility to have high confidence in future revenue growth and to invest as needed ahead of it. Our existing large strategic customers provide validation, contractual visibility, and the economic support required to build new capacity at scale.

Speaker 3

At the same time, we are having success expanding our addressable market and customer base. OpenAI provides, among other things, scale and frontier insight. AWS provides global enterprise reach. Recent collaboration with AMD expands the market opportunity to include disaggregated inference, and fast inference is cracking open more new markets like security. We ended Q2 with more than $8.6 billion in cash equivalents, restricted cash, and marketable securities. We also have a revolving credit facility of up to $850 million that has been unused to date. Our liquidity and balance sheet position is strong and was enhanced by our IPO in Q2. It is a significant advantage that provides us with flexibility to invest and adjust to opportunities in these very dynamic and high-growth market conditions.

Speaker 3

In addition, we have the advantage of much lower net capital expenditures per megawatt than the vast majority of AI cloud providers for two key reasons. First, we primarily incur CapEx for the deployment of our own hardware and our data centers at much lower BOM cost that does not include the high profit margins many others must pay. Second, we are reimbursed for a meaningful portion of the remaining CapEx for data center fit-out as data center pass-through cost reimbursement from our largest customer. Now turning to our outlook. For Q3 2026, we expect core revenue to be in the range of $214 million-$216 million, core gross margin in the range of 38%-40%, and core operating margin in the range of -25% to -23%. For the full year 2026, we are raising core revenue to the range of $880 million-$890 million.

Speaker 3

We are raising core gross margin to the range of 41%-43%, and we are raising core operating margin to the range of -19% to -17%. In closing, Q2 was a very strong quarter of continued execution and growth for Cerebras. We delivered record core revenue, cloud and services revenue nearly quadrupled, and gross margin and operating margin were also significantly better than our guidance. We improved our guidance for each of these items for the full year. We have made great progress building our capabilities, capacity, and customers, and ended the quarter with more than $8.6 billion in cash and cash equivalents and investments to continue to execute our growth plans. We are well-positioned to grow revenue by more than 3x in 2027, and for tremendous additional growth in the following years, while also significantly expanding gross and operating margins towards our targets.

Speaker 3

I'll now turn this over to Andrew for final thoughts.

Speaker 2

Thank you, Bob. More than 10 years ago, we started Cerebras with the belief that we could build a better processor for AI, and the belief that to deliver the processor, we would need to build a full accelerator system and racks. Today, Cerebras is one of only four companies, Google, Amazon, NVIDIA, and Cerebras, to build processors, systems, data centers, and deliver AI-based cloud services to customers. Thank you for listening to our prepared remarks. With that, I ask the operator to please open the line for questions.

Operator

To ask a question, please press star one one on your telephone and wait for your name to be announced. To withdraw your question, please press star one one again. In the interest of time, we ask that you please limit yourself to one question and one follow-up. Please stand by while we compile the Q&A roster. Our first question comes from Timothy Arcuri with UBS. Your line is open.

Speaker 4

Thanks a lot. Andrew, I wanted to ask about customer concentration. You did say that revenue would be up more than 3x next year, and obviously we know that OpenAI is ramping right now, so that's a big piece of your incremental revenue today. I would think that AWS could be $1 billion next year, something like that, maybe more. So how do you think about customer concentration when you look at next year? Is it going to be two-thirds of your revenue is those two customers? Can you also speak to your talks with some of the other folks, Google and Microsoft and folks like that? Thanks a lot.

Speaker 2

Sure. I think that's a good question, and I think some historical perspective might be worthwhile, right? In 2021, people complained that we only had government customers. Then when we won a sovereign cloud at $1 billion, there were concerns we only had a sovereign cloud. Then we won the largest frontier lab, and then there were concerns that we didn't have a hyperscaler, and then we won AWS. I think in each of those cases, we were able to use the momentum that the previous step gave us to expand our business. I think OpenAI is an enormous customer, and they're an enormous part of not just our business, but of everybody's business in the sector. I think they'll stay a big part next year.

Speaker 2

But you're absolutely right that AWS and others, whether they're the rapidly growing coding companies or some of the use cases around security, they will be a larger portion, and OpenAI will shrink as a percentage of our revenue over time. But I think you can expect for next year them to still be a meaningful portion of our revenue.

Speaker 4

Great. Then just as a quick follow-up. I know, Bob, you said that capacity, I think you said it's going up 10x this year-over-year. Is there any sense of how much it's going to grow next year? I know that Andrew said that revenue's going to grow 3x, but is there any sense in terms of how much your manufacturing capacity will actually grow next year-over-year? Thanks.

Speaker 2

Yeah. We're going to end the year well over 10x our manufacturing capacity. We already have contracted facilities three or four times more for growth in 2027, and we still have some time to contract for more. We're looking at enormous growth over a several year period.

Operator

Thank you. Our next question comes from Joshua Buchalter with TD Cowen. Your line is open.

Speaker 5

Hey, guys. Thank you for taking my question. Following up on Tim's previous one, can you maybe just walk us through how the economics of the Amazon deal are going to work? Is the plan that it will be available next year and offered in AWS cloud, and then we will basically see how much demand is and so it is difficult to forecast right now? Thank you.

Speaker 2

Sure. It will be available. It is deployed in Amazon data centers. It will be delivered through Amazon's API service, Bedrock. We are in the process right now of organizing deployments. I think that is sort of the way to think about it. We expect the service to be live in Q1.

Speaker 5

Got it. Thank you for that, Andrew. With the AMD engagement, any more color you can give on the go-to-market as you connect with the Helios rack, and timeline you would expect to revenue. Regarding the AMD engagement, they made an acquisition of an inferencing hardware company recently. Could you maybe speak to how that fits in with what you guys are offering as we think about their broader suite? Thank you.

Speaker 2

Sure. I think a couple things. I think that we will be announcing additional parts of our arrangement with AMD over time. I think the joint solution of Helios racks in front of Cerebras Systems, the Helios racks doing pre-fill and Cerebras doing decode, is an extremely strong offering, right? An offering in which we deliver vastly faster speed than Helios can deliver and vastly more throughput than Cerebras can deliver alone. That solution is enormously compelling, and we have buyers for it already. The second question is of recent acquisition by AMD. Look, I think the company they acquired was interesting and innovative, and no one is more excited about innovative hardware than we are. I think buying hardware startups, there is a lot of time between when you buy them and when they deliver.

Speaker 2

I think we were impressed by what those guys were working on, and think there are many applications in AMD's portfolio for them. I do not see the first application there being data center inference.

Operator

Thank you. Our next question comes from Tom O'Malley with Barclays. Your line is open.

Speaker 6

Hey, guys. This is Kyle Buser on for Tom O'Malley. Thank you for taking our question. I wanted to go back to Josh's question on the economics with the AMD deal. Is the way this kind of works, you buy an AMD Helios rack, install it in your cloud, and then all the revenue that comes from customers renting out this disaggregated inference solution goes to you, or is there some sort of revenue sharing agreement that happens here?

Speaker 2

Yes to the first part of the question.

Speaker 6

Okay. Thank you. For my follow-up, the AWS deal is getting installed in their clouds first. Do you see an eventual path to you hosting Trainium and the CS-3 together in your cloud or another hyperscale cloud? Just trying to think about how disaggregated inference can evolve in terms of future deployments.

Speaker 2

I think we are very interested in that approach. I think that, as you've seen with Google, there is an opportunity for hyperscalers with their own parts to seek to deploy those parts outside of the boundaries of their own data centers. That's something we'd be interested in, not just with AWS, but with others. I think that's very much on the table for the future with AWS.

Operator

Thank you. Our next question comes from Quinn Bolton with Needham & Company. Your line is open.

Speaker 7

Hey, just wanted to just a quick clarification on the AMD deal. Andrew, if you purchase and stand up the Helios racks in your Cerebras cloud, but that service is delivered to OpenAI under your contract, does that represent an additional revenue opportunity, or how should we think about the potential for revenue in that instance? I have a follow-up.

Speaker 2

I think that whenever you increase throughput while keeping your performance the same, you increase your opportunity for revenue. Throughput is the number of customers that you can simultaneously support. If you can do that without giving up speed, you have more revenue per system. I do not want to go into the specifics of our relationship with OpenAI, but one of the things that makes us so excited about this partnership is that you keep our speed and you increase throughput, which makes each system more profitable. Each system is generating more Tokens. That means tokens cost less, tokens use less power. Not only do we make more on top line, but our margins improve.

Speaker 2

Not only do our margins improve and our top line improve, but it makes each data center investment more valuable, because a data center is a power envelope, and if you can get more tokens out of that power envelope, that converts to more dollars. It is an enormously powerful thing and with AWS doing disaggregation with us and with AMD doing disaggregation with us, we have tied up about half the leading chip makers. It is a very, very powerful story.

Speaker 7

Very fun. The following question is just, you mentioned in the script a couple of times that you will increase your throughput of the Wafer-Scale Engine by a factor of 20 by the end of 2027.

Speaker 2

Correct.

Speaker 7

Does that remove the need for some of this disaggregated compute, or heterogeneous inferencing that you're talking about? Or does that just make the entire throughput of the heterogeneous solutions just that much faster?

Speaker 2

I think we're

Speaker 7

just a faster, higher throughput.

Speaker 2

I think we're exploring all sorts of ways to drive throughput up, right? If your throughput increase is 20x and your costs stay the same, you're in pretty darn good shape, right? So our systems are improving throughput. We're looking for ways to improve the throughput of disaggregated solutions. We're looking at all sorts of different inventions, technologies, partnerships that continue our sort of pattern of industry-leading performance and vastly increasing throughput.

Speaker 7

Thank you.

Speaker 2

That's a really good question. That is what we're thinking about in our roadmap, that exact point.

Operator

Thank you. Our next question comes from Joe Moore with Morgan Stanley. Your line is open.

Speaker 8

Yeah, thank you. On the lines of what you were just talking about, when you talk about disaggregated decode, where are you in terms of commercialization of that? We know you can do fast inference at scale. You've done it. When it comes to disaggregation, is that ready to deploy now? And what work needs to be done over the next kind of year to get to the types of improvements that you guys are talking about?

Speaker 2

Sure. I think we have disaggregated inference with GPUs running in our labs right now. I think it will be deployed and available in Q4.

Speaker 8

Okay. Thank you. You talked in your script about the ability to work with the installed base of GPUs. When you work closely with Amazon, work closely with AMD, is that stuff going to work better than what you would be able to do with like NVIDIA installed base GPUs that are out there?

Speaker 2

I think it's fair to say, though we haven't done it yet. I think it's certainly fair to say that Helios racks will give us a bigger and better solution than if we were to use the B200s, right? Trainium3 will give us better solution than if we were to use Trainium2. If we were to use other GPUs that the current generation, the top of tree generation, will give us better performance than the top of tree minus one generation. But I think it's also fair to say that the minus one generation will be vastly better in a disaggregated solution than not in a disaggregated solution. All right? In an environment where everyone's trying to extend the life of their hardware and continue to keep it delivering valuable tokens, this is an important option. Does that make sense?

Operator

Thank you. Our next question comes from Vijay Rakesh with Mizuho. Your line is open.

Speaker 2

Hey, Vijay.

Speaker 9

Yeah. Hi. Hey, Andrew. Just a couple of quick questions. You mentioned the 600-megawatt signed capacity, signed power, and then 10x increase in capacity by end of the year. Do you think that should help you accelerate some ramps in 2027?

Speaker 2

Yeah, of course. I think that we are pursuing data center capacity around the world every day, and we're doing it because we have tremendous demand for fast inference. The faster we can deploy, the faster revenue grows. That's true not just for our cloud business, but it turns out to be true for our customers' on-prem business, right? The faster they can get data centers, the faster we can ship them hardware. It is top of mind. It is something I spend an enormous amount of time on. We have a whole team now. We're pretty darn good at chasing down data centers around the world. Once you've signed them, your job isn't done, Vijay, as you well know. We have people on site every day.

Speaker 2

We are engaged with the developer and the construction firms at every stage to do our best to keep them on track. The faster we can do that, I think the faster we can ramp our revenue.

Speaker 9

Got it. As you look at partnering, I know you mentioned hyperscalers, but there's a whole emerging neo cloud group that's coming up. They're getting financing. There's a lot of financing structures being developed across Wall Street, I guess. How is that pipeline developing for you? Thanks.

Speaker 2

Sure. I think early on, the neo clouds were very focused on NVIDIA. I think as the business has become clearer to investors, there are neo clouds that are diversifying and find themselves less dependent on one hardware vendor. The opportunities for us in that category are large. There are neo clouds that are multi-vendor. There are neo clouds that are AMD only. There are neo clouds that are coming up out of people who have power assets. I think in 2027, they'll be an important part of our business.

Operator

Thank you. I'm showing no further questions at this time. This concludes today's conference call. Thank you for participating. You may now disconnect.