Forward Guidance: AI Efficiency Is Repricing The Compute Market | Steve Hou
AI’s next phase hinges on a paradox: falling costs could threaten today’s winners while unlocking far greater demand. Steve Hou, head of research at Silicon Data and former Bloomberg strategist, join
view source ↗Forward Guidance: AI Efficiency Is Repricing The Compute Market | Steve Hou
Sourced by
podcast-ingeston 2026-07-22. Auto-transcribed via AssemblyAI (universal-2,en). Speakers identified by AssemblyAI Speaker Identification using the per-podcasthost/regularshints; the resulting label→name mapping is in the frontmatter. Duration: 43m. Episode page: (not provided). Audio: https://traffic.megaphone.fm/BWG4187305663.mp3.
Show notes (from RSS)
AI’s next phase hinges on a paradox: falling costs could threaten today’s winners while unlocking far greater demand.
Steve Hou, head of research at Silicon Data and former Bloomberg strategist, joins us to examine the changing economics of AI compute.
We discuss token efficiency, model routing, GPU pricing, memory bottlenecks, and when enterprise adoption may finally deliver measurable returns. Enjoy!
TIMESTAMPS:
00:00 Intro
01:01 Why AI Compute Needs Hedging
06:55 What The Token Index Really Shows
14:04 Token Maxing Meets Efficiency
18:35 Who Captures AI’s Value?
22:12 Old GPUs Reveal Surging Demand
27:10 GPU Markets Keep Tightening
32:03 The Memory Bottleneck
37:07 AI’s Next Phase
FOLLOW STEVE
› X/Twitter – https://x.com/stevehou
› Silicon Data – https://www.silicondata.com/
FOLLOW THE SHOW
› Forward Guidance – https://x.com/ForwardGuidance
› Felix – https://x.com/fejau_inc
› Telegram – https://t.me/+CAoZQpC-i6BjYTEx
› Blockworks – https://x.com/Blockworks
EVENTS
› Join us at Digital Asset Summit 2026 Asia October 7th & Digital Asset 2026 London November 10-11th
DISCLAIMER
Nothing said on Forward Guidance is a recommendation to buy or sell securities or tokens. This podcast is for informational purposes only. Any views expressed are opinions, not financial advice. Hosts and guests may hold positions in the companies, funds, or projects discussed.
Transcript
Jack Farley: Nothing said on Ford Guidance is a recommendation to buy or sell any investments or products. All right everybody, welcome back to another episode of Forward Guidance. And joining me today is repeat guest of the show, Steve, who who just joined me a couple months ago right at the the tail end of when you're at Bloomberg but now you're at Silicon Data, head of research. You are the man behind some of the most important charts and indices in the world of AI right now. There's a lot to get into, but yeah, really excited to have you back now under your new role and your new company. So congrats on starting there and yeah, great to have you back, Steve.
Steve Hou: Yeah, thank you, Felix. Great to see you always. And thank you for having me back on.
Jack Farley: Yeah, awesome. For those that don't know about Silicon Data and what you guys do would just love to hear a bit about like why you made the shift to joining them, what you guys do and yeah, just how you think about the state of AI right now after.
Steve Hou: So yeah, I left Bloomberg about two months ago, joined Silicon Data almost two months ago back on the day. And Silicon Data is a company that tries to bring data to physical AI compute market. And I think the easiest thing to think about is we are trying to bring futures contracts, derivatives to the AI physical compute and allow people to hedge the essential risks that are involved with this now enormous and evolving AI compute market that's hundreds of billions if not trillions in size. And there's a lot of risks that's being held in equity form and in a fixed income form and in our view sometimes probably not perfectly efficiently. And there's a lot of I think risks that's now involved with AI compute and GPU income data can be handled with traditional financial instruments that is well understood by historical financial markets or capital markets. So to the extent that on the natural hedging side of the designer providers, computer providers or the companies that are looking to buy compute rank compute futures contracts is a very natural way to hedge out their risk and maybe actually in fact help you be a bit bolder in terms of at the outset how much computer actually acquire. So you don't find yourself being, I think overexposed or maybe not having enough, which seems to have been the case with some of the AI labs. I joined the company because I felt like I have been working on public markets benchmarking systematic indices at Bloomberg for six years. This feels like a natural way to become exposed to the AI sector. Same sort of similar set of skill set but apply to AI.
Jack Farley: Yeah, awesome. Yeah, so why don't we get into, we'd love to just hear a bit about like, who, what is, what is the. Why is hedging and you know, futures contracts a necessity for the AI compute build out, like who are, who are the ideal target customers and what are they trying to hedge and why, like, is it the hyperscalers, is it the frontier models? Is it like you know, just downstream companies that are utilizing these models and you know, are trying to hedge out their inference demands or like what does, yeah, what does that landscape look like?
Steve Hou: So, I mean, the demand for sure really come from all major participants in this market in the same way you would expect in the crude market or even agriculture historically, wherever futures contracts first evolved, as this market eventually become more, I think, fragmented.
Jack Farley: Right.
Steve Hou: Right now the AI compute market is very much dominated by a couple of big players, a couple of big sellers, very, very granular. Right. So you've got the two major labs, OpenAI and Anthropic, that account for ostensibly half of the AI compute demand and then on the other hand, the hyperscalers that are providing most of the compute and building most of the compute, as we have now seen with the rise of very powerful open models and enterprise AI use. Actually if you think about how enterprises will ultimately adopt AI, they're going to actually not just basically pick a winner, debate of who is going to emerge as the sort of platform on which everyone's going to build on top of, I think that's already been settled, right. Nobody's going to win it outright and companies don't feel comfortable and the labs themselves, I think it's just not. So what's going to happen is that orchestration layer is going to accrue a lot of the value. Companies are going to retain their sovereignty by basically making models more substitutable and some tasks, low value tasks maybe going to be routed smartly towards the open models, cheaper models, and then the high value tasks maybe would be using frontier intelligence. To that extent. We expect there to be generally a transition away from the current paradigm, even more so. And with inference becoming an ever bigger deal that most of the AI computing is going to come from inference and the funding of compute is going to come from many, many enterprises, companies that will be, I think, looking for compute. And as the models evolve, the architecture of the models evolve, you're going to see that more models can be run on not just the hyperscalers, but may be sort of more closely located, smaller clusters, smaller clouds. So you can see this fragmentation. So the market on the one hand you have sellers of compute, these new clouds, cloud providers coming under the market. They want to have more certainty and their lenders, their backers want to have more certainty over their revenue. Right. And how do you actually have certainty? One natural way is to actually just as you have with any commodities market, using futures contracts to lock in future revenue. On the other hand, you know, you have the buyers of COMPUTE and different ways in which you can work out. Some companies, maybe they don't directly buy compute, maybe they just buy tokens, but somebody will actually be sort of buying that compute. Right. So I think that's how we sort of see it eventually the market evolving that direction. Now of course, along the way, you know, if you talk to cloud providers today, there's going to be a lot of pushback. Nobody wants to be their product to be referred to as a commodity and. But I think the way I see it is that there's always going to be substitutability. Right. If you tell me that your product is so unique, how do you win customers from someone else? How do you win clients away from someone that doesn't. With someone that. So there is going to be substitutability when it comes to compute, but it's not going to be perfect. And that's how we see this market.
Jack Farley: Awesome. Okay, so I want to dig into a few of the different charts and indices that you guys have built out and want to start with just on that point about orchestration and substitution. Want to talk about the token index that you guys built out. This one has gone sort of viral for maybe good reasons, also bad reasons. There's been a lot of takes on it. I would make the argument some quite uninformed take. So I just want to like let you explain through the methodology, just from first principles. Like I'm not going to throw any sort of bias here. Just this is, this is your token index. Tell me about the methodology of what attracts and what we can take away from this chart.
Steve Hou: Indeed. Yeah, indeed. This, this index was sitting on the shelves, you know, when I joined the company and you know, for quite a couple months and you know, and I think it was not getting a lot of attention. I think I just noticed that people were sort of. Whoever noticed it was actually not interpreted even correctly to. I thought it was either price index or total volume index. I think the name of the index on terminal displayed is also got a little bit unfortunate. There's a bit of a misnomer because it's called the token expenditure index. When it's actually neither. What it really is is an expenditure weighted price index. You can think of as almost like the PCE for AI, right? So what that means is that to the extent you're on track inflation of something, you can either fix the basket, which is what CPI does, or pce, you allow substitution across similar substitutes because users are rational, they always make the quality price trade off and AI even more so. Especially when you have sort of these open routing or inference platforms that allow you to choose different models, right? You know some models are really, really powerful, but they are crazy expensive, right? You know your fables of the world, right that burden half of your token budget, you know, in a day or something. But then you've got the Chinese open weights models that are super, super cheap, but they are maybe less capable. So I wrote a post on X on our Corporal X account towards the beginning of June saying that this is the way you should be interpreting this index. We basically collect all the prices of all the models, 300 plus models out there. And also we collect a usage volume from a couple, a few sort of these public routing platforms with the input output volume. So as you know that every ALM token model sort of there's a price charge on input token how much tokens you put in and there's also a price charge on output tokens. So we have to create a blended price for every single model and we sort of take the input price, input volume and output volume, create a blender price beige model and aggregate them up to index level based on usage so that this thing is going to track how much each model is being used. And to the extent that the general price dynamics of this model is such that once the model is launched the token price moves, but not that much, right? So most of the movement comes from essentially usage pattern, usage behavior, consumer behavior. So the, our coverage of data is not a full market. If you imagine the full market, the AI inference market as being like a square, this is not a vertical slice of it, it's not a representative sample because we don't have data to. We're not privileged the data of say OpenAI anthropic internal data usage or the hyperscalers for the matter. We are talking to a couple of them. But at the moment we sliced off a corner and this corner tilts admittedly towards probably like independent developers or small medium enterprises that are already onto some of these inference platforms, right? And they are going to be necessarily the ones that are more price sensitive than you would imagine, like a large enterprise that's currently living off of a hyperscaler that maybe the employees are not even responsible for the expense and they just go crazy until like end of the quarter, end of the half year. And the CFO say wait a second, like so this is going to be a lead indicator. And I wrote a post at the beginning of June saying this thing seems to be plateauing a little bit and the use correct situation is such and this thing could be taking a break before links leg up. Maybe we see a super powerful model and anthropic just runs away with it. Or there's a possibility that this thing actually mean reverts a bit as people become more rational about the cost of these things and sort of substitute between quality and cost. Indeed that seems to have been what happened and also unfortunately coincided with the turn of the stock market. Because I think partly rationally because the current paradigm of how the AI Capex trade has been funded and everything's being priced on second derivative. So even if things keep going up just like a slower pace, things are going to sell off just based on the valuation. So there's that. But the fundamental is such that if there is a perceived slowdown, a rationalization and there's a perception that the frontier models could see their margins challenged with very powerful much much che cheaper open waste models or even just sort of not open but capable but cheaper models coming from Groq and Spark and whatever that this could actually challenge the current paradigm of how the air capex is financed. And I think this is part of the reason why we've seen markets selling off. If you open overlay this chart with the AI sort of beneficiaries index or the SMH or whatever, you can probably see that turning point. But we don't necessarily see this as like a secularly bearish index. I don't think there is actually somehow contradiction and this is not one of those things where it goes up is bullish, goes down is bearish bullish either. It's actually a little bit conditional, right? Just like you would never say oh, higher inflation is always good or lower inflation is always good. At some point we're going to transition towards that paradigm I was talking about earlier where the funding of AI and the demand of AI becomes a lot more broad based and you can have a lot more sort of smart routing of models even as the inference market explodes in size become a much, much bigger of the overall pie of AI demand and AI compute. So this transition from one phase one to another was always going to be a bit rocky and even potentially runs in the risk of a bit of an air pocket where on the one hand the capex and ROI for the current paradigm doesn't show up quickly quite enough. On the other hand, the new phase, maybe the biggest enterprises, takes them a while to find that workflow and pick up. So the pace may not catch up. You will see the fastest adoption and most ROI show up in the smallest enterprises with the fewest employees and the least frictional workflows. But those are going to be by dollar amount, by dollar weighted, going to be small. So part of the reason why we don't see productivity increasing in aggregate part is also because of that. Right. So anyway, that's kind of a long winded answer to. I don't even remember what your answer a question was, but hopefully still made some sense.
Jack Farley: Yeah, no, that's super helpful. I think there's some, some really important nuances in there. But yeah, I mean, I'll throw the proposition at you. That was those before, which is that. Okay, if you look at the chart, you know, early March, which is around where when Mythos started to come out, it was also around the same time that this, this idea of token maxing took hold of of U.S. corporate America. And suddenly it's just like, okay, just token maxing use as much as possible, use the frontiers. And then that chart went higher and then I think what some folks took away from. And I think to your point, you mentioned how partly it was just this index, maybe it was a bit of a misnomer, but people took it as okay, this is just actual demand for tokens from corporations. And then you see the turn came forth right around the time where these news headlines started to come out of like Uber went through their entire year budget of token spend in like a month and everybody freaked out. But what you're saying is it's more so about substitution than outright token demand, right? Like what you're saying is that this is about going from token max into token efficiency. It's not saying that demand for tokens has gone lower since June. So we just love. Yeah. Can you expand on that further?
Steve Hou: Yeah, absolutely. I think this idea of token maxing was always the idea you want to actually let people experiment, right. Because complex workflows, large enterprises, you don't really know what the use case is going to be. And also the models just didn't really become smart enough or good enough until this year with agentic and everything. So you need Token magazine. But Token magazine is fundamentally odds with expensive tokens. Tokens are so expensive and also throttled. So you need eventually to get to a state where tokens are cheaper. And this idea of smartly routing models that's always going to be the case if you look at my old employer Bloomberg, this very nice paper on MCP of how they let the models interact with the data and correctly quality accurately with financially that you can't really afford hallucination. You see what's been happening at Palantir or this idea of ontology of actually making the AI actually useful at the enterprise context and databricks same thing. And more and more of this open routing has. Has turned from sort of open router will give you this sort of routing for as long as time. That seems to be more of like a hobbyist developers trying to try different models. But I think that paradigm has actually become now the norm. Right. Increasingly we should expect that. Right. And this is kind of like to the extent I hate to use the term now because it's so cliche but like sort of the Javon's paradox Max basically in the fuller context of it ultimately how this thing can be truly sustainable and actually get an ROI is to actually get enterprise adoption. Millions of companies and the only way that you can actually get sustainable adoption of AI is when you don't suck dry the underlying company. We just have a single model the entire company is built on top built on top of. And that was not going to happen. It certainly is not going to happen now. Instead basically the model layer has got to be substitutable where you can smartly substitute between. If I have a simple task, say just data cleaning or just choosing what analytics or simple summary summarization or whatever things that can be easily handled with much cheaper less than frontier models, they should go to those things. And if I have a higher value to ask certain things that requires more powerful reasoning, they should also go to those. And the ability to route between them and correctly assess quality and cost is something that will have to happen. And it's not really been happening enough and is a challenging thing because just to figure out whether or not a certain ask is which type it's not always obvious. Sometimes a question can seem benignly innocuously simple, but in reality it's actually a pretty complex question that someone is just asked poorly because they don't really know what they are looking for. And I think that is going to be the new challenge. And that's part of the reason why you are seeing more and more of these sort of routing things being activated sort of vercels of the world or ramp is coming up with one and so on and so forth. So this is going to be the new norm.
Jack Farley: I think that's super helpful. All right, Steve, so you're an economist, so I want to ask you about this. But it feels like one of the big questions right now that sort of leads into what we just talked about with this index is this idea of feels to me like this framework is either the frontier models get to enjoy the margin. Like right around this time where this chart was, was surging was also the time where the ARR of anthropic was surging. And you know, Gavin Baker mentioned that there was even a month where anthropic was like, was profitable for the first time. So that's all great. But now we're seeing this substitution towards token efficiency. We're starting to see these open way models that are looking a lot more powerful. We're starting to move to this world of perhaps frontier model top line orchestration. But a lot of the downstream execution is these more efficient, cheaper open weight models. It feels like that is really good for the end consumer. The surplus goes towards them. So it feels like there's this fight for margin between the frontier models in that AR and then being profitable and everything that is almost dependent on that occurring versus the benefit going towards the consumer. So I'm just curious, what's your framework for how to how that balances out? Like can the both sides win or what do you think?
Steve Hou: Yeah, I mean this is exactly like you asked me as an economist, right? You know, economists like to think about partial equilibrium and general equilibrium, right? So what is partial equilibrium? Partial equilibrium is when you just have like a simple shock to price, right? Price times minus cost times quantity gives you profit, right? So cost is kind of whatever, you know, cost of compute and operating. And if price is challenged by a competitor, your and quantity doesn't change. You're going to see a shrinkage of the overall profit. And that makes you worried about sort of ROI and so on and so forth. But Q is not static, right? That's the whole point of Jimin's paradox is there is elasticity, right? And if the overall market has become much bigger right now, the adoption, enterprise adoption is abysmally small. Most people are not using AI in a consumer context beyond sort of glorified search. And enterprises are still trying to figure it out and they are nervous about the cost costs even as they just figure it out. Now if you tell me that, okay, the margin of maybe some of the frontier Models going to get challenged a little bit, but the Q is going to grow so much, much more. The overall thing still works out. Right. You know, the analogy I like to give a little bit sometimes it's a bit like this is not totally unlike the drug market. Right. So people always ask who is going to innovate on pharmaceuticals if you just get generic drugs that sort of copy you and whatever. So there's going to be some dynamic of, you know, erosion there. Right. You know, you're probably not going to earn like full and monopolist profits, you know, if you have competitors at the same time, I also don't think it's been true. Right. Whether you look at software or look at pharmaceutical or whatever, just because there is going to be competition and there's going to be some degree of copycats or what have you, it doesn't mean that the frontier is not going to grow and it doesn't mean that, you know, the leaders of the frontier will not be profitable. That being said, I do think this is going to challenge the question where the ultimate bulk of the value is going to accrue to. And I don't think it's going to be super controversial to think that at least direction on the margin. This means that some of the value that was perceived to have been certainly accrued to this duopolis model of frontier intelligence labs, probably more accrued to either the compute layer or the user layer, depending on how the compute layer settles out. Right.
Jack Farley: So makes sense. All right, that's a good segue into talking about neoclouds and GPU rental index. So this is another index from, from you and your team. Walk us through how to think about this. And, and yeah, what does this, what does it say for the current landscape?
Steve Hou: Yeah, so basically we track, you know, create these, you know, rental indices. It's almost like a benchmark. Can you think of analogous. Like if I want to create an unfurnished single bedroom apartment index in New York City, you observe rental contracts from across the city, a different term in different shape and form. Some come furnished, some come unfurnished, some with a little bit of room and so on and so forth. We aggregate them all. Use machine learning to essentially create apples to apples comparison and create a singular index for each chip. Out of all the data we look at. And here we have the Blue line, that's H100 chip hopper. Right. And then you've got a 100, which is the engine chip that's been out for like five years or something, a 100. And then you've got the B200, which is the orange line is more volatile and also newer, less deployed. And also the H200, we have a longer history, but here only for some reason, I only managed the rent of the most recent few weeks. So the way I look at it is that, you know, amateur who follow the AI market, they just look at the H100, you know, sort of index, which is really actually the on demand price, right? This is not the entire four curve, right? This is not telling you like what how much will cost per GPU hour out, you know, year out or whatever. And they look at this blue line, if it goes up, they say, oh, you know, demand is going up, it's going down. Like, oh, demand must be collapsing. But in reality I always like to both look in the cross section and then later, you know, we'll get time to also look at across the term structure. Because in the cross section what happens is that between inference, demand and training, especially with inference growing so much, you can actually see the shifting of workflows and depending on different labs and sort of who's needing seeing, demanding what kind of thing, you can actually see shifts. So for example, when H100 has been steadily going up since the beginning of this year and it has had some, you know, sort of fluctuations recently came down. When it came down, what you see is that the purple line and the yellow line or the orange line continue to go up a lot, right? So meanwhile, if you look at the A100, you know, that's actually very steadily sort of holding, right? It's not even coming down at all. You can have a scenario which, which, you know, this is consistent. I'm not saying this, no, for this for sure. But you have a scenario where inference demand is just through the roof, roof growing like crazy, right. And that's driving up rental rate for even a 100 chip, like you know, from five years ago and strong. Right. Meanwhile, you can have some of the workflow for training shifting away from H100 towards the newer chips, right. And, and as those chips come online and, and depend on availability, so those things can fluctuate a little bit. So that's how I read the, the at least that on demand, you know, in the indices for.
Jack Farley: Yeah, awesome. I feel like people look really closely at this because of course one of the largest companies in the world right now is Nvidia. So obviously demand for the GPUs is very correlated with the performance of NASDAQ and the total stock market. So obviously if there's any Sort of concern for demand for GPUs in the build out that could have some pretty significant shockwaves throughout the system. So is there anything that we can take away from what we're seeing right now in this indices? Interpolate where GPU demand is at and whether the build out is peaking out or whether the growth rates are changing there. What do you see there?
Steve Hou: So we can get an even cleaner read from the forward curve, which I think we hope we can come to next. But I think on this chart, the favorite thing I look at on this chart is actually the green line and not so much the blue line. Just look at how strong that price is, people. There was all this talk about early in the year, like Michael Burry, like, oh, chips are only good for two, three years, the fast depreciating asset, blah, blah, blah. In reality, even the a 100 chip rental rate is going up and has stayed up, has not come down at all. Right. That just tells you just how robust inference demand is. And going back to what we were saying earlier about routing different workflows and whatever, there's just a lot of low hanging fruit, easier tasks, inference, demand. As this overall demand grows, people become familiar with AI. A lot of those things are being routed towards the workhorse chips. So I wouldn't be surprised if all the most widely deployed hopper chips, which will soon become no longer the most powerful chip for training, will actually become a new workhorse. So you will see sort of a similar pattern to follow here. So that's actually a good signal. But then we can also look at the forward curve from.
Jack Farley: Right, yeah, walk us through there and how to understand it.
Steve Hou: And yeah, so basically focus is the way I think of is the following, right? So if you are running a GPU like, you know, some of these contracts can run for a fairly long time, right? You know, you can run it for a month, you can run it for three months, you can run it for as long as two, three years. And typically with any commodities curve, you expect a downward sloping like, you know, backwardation or, you know, that's the jargon, right? So downward sloping curve of like, you know, because we expect supply to come online in this case, like another example you can imagine is like if you rent an apartment in New York or anywhere for 12 months, your per day rental rate is probably going to be lower than if you rented a hotel room for three days, right?
Jack Farley: Why?
Steve Hou: Because you get, you get a discount for avoiding the trouble for the landlord, the turnover, tenancy, right. You sort of lock in that thing and so it doesn't have to sort of keep looking for the next renter. And indeed that's what we saw like the blue line here from November 22nd and last year, you know, just when the agentic AI was sort of kicking off the opus models becoming online. The, the, the, the curve was very backward data, downward sloping. And since then like between November to like this is a lot of curves. I want to show you sort of how things have been evolved in the last few weeks. It's interesting, right? The, the orange line here is the. At the end of March, right, March 31, right. The entire curve essentially moved upward. Right. In other words, every maturity, every length of the contract per GPU hour price has gone up. Not only that, the curve has also become on the long end almost seemingly a little bit in contango, right. Like sort of no longer downward sloping. What that means is that cloud providers, renters, the landlords of AI clouds, are feeling comfortable not giving these long term contract discounts and just letting short term contracts roll over so that they have an opportunity to raise prices again. And indeed that's what been hearing from anecdotal evidence with anecdotal conversation we have had with cloud providers that they really are very much aggressively looking for opportunities to raise prices. Now some still want that certainty. So this is why there's a market. Earlier we were talking about hedging. People have different views. If you have a thing, maybe you grow some crops, maybe you throw some oil, if you feel like prices can just go straight up, you don't hedge a at all. So what's interesting here is that between March and June, if you look at the purple line, that's like June 25th, there was a bit of a scare in the market capacity basically on the front end. H100 came down a little bit and people freak out a little bit. But then you look at the one year mark, there's a lot of the meaningful compute workflows that actually contracts that like the one year term that's pretty much unchanged. And what's even more interesting is that in the ensuing weeks over the last month, you have all this news about meta leasing out compute and all these models and substituting away from the frontier models. Basically a bunch of news that will be indicative of AI supply maybe potentially being in access because the market valuation is so high, people are looking for every reason to freak out. What we see in the fundamentals, at least in the terms of compute, is that rental prices at the one year mark for the fall rate, it's actually been monotonically going up. There are two lines that overlap each other because there's not much change between 2nd of July and 14th of July. And then most recently when we look at yesterday, July 22nd 20th, the Green Line went up even more every single time we saw multiple of all providers of actually raising prices. In other words, it's not just like a one off thing, at least in terms of compute fundamentals. There is actually pretty sort of strong indication of firming demand and supply being in shortage. Therefore the price has to adjust.
Jack Farley: Awesome. Anything you want to add on the other GPU curves here in terms of, of distinction?
Steve Hou: No, I mean I think pretty much similar, you know, like, I think you see a similar picture on the right, you know, sort of there's the A100 and is you know, sort of secularly going up, right. And you know, maybe not quite as aggressively, you know, more recent weeks, but generally holding up pretty strongly. And then the B200 is also generally moving up, but then there's a bit of a fluctuation. So the thing with B200 and also I think just think about this chip dynamics generally is that B200 is still being deployed, right? Most of the data centers are still just bringing them online. You know, availability is not quite as quite as high as the uh, 100, so there's got to be some, you know, sort of fluctuations, you know. But generally speaking, you know, I think the picture that emerges qualitatively is very much in agreement totally.
Jack Farley: Obviously we've been Talking mostly about GPUs, but increasingly a lot of where the, the change in, you know, the build out for, for all this AI is also on the, the DRAM side of things. And the memory aspect, I'm curious how you think about that and what's going on in that side of the market compared to the GPUs.
Steve Hou: I mean so clearly like some memory is input right into GPUs and all these models are super memory hungry. And generally speaking with longer context, longer conversations, the context grows with memory. That's part of the reason, that's clearly the reason why memory, as far as storage, you generate tons and tons of data that all just demand is not catching up with supply is not catching up with demand. And it's not surprising that price is shooting up the way it is. But we also know that shortage and high prices are always the mother of innovation. So unsurprisingly we're seeing algorithmic innovation no less from the recent Chinese models like Kimi and so on. That are actually making sort of improvements to memory efficiency and sort of maybe the memory demand doesn't have to quite grow linearly with context length, which seems to be, I think, part of this KDA innovation that came with the community model. I think it's not surprising to me the Kimmy thing moment felt very similar to the deep Seq moment. Like we're going to keep on having, I think efficiency improvements like that. I mean, clearly it's not going to be like I would be very, very shocked if, let's say in two years the gross margin of the memory makers are still running at like 85%, you know, whatever. It's just like one of those things is a sort of a squeeze.
Jack Farley: Right.
Steve Hou: So I think, you know, we get past it, you know, and you know, I think there will be again goes back to the question of P times Q. Right. Sort of. I think P is going to come down some, but the Q probably grows enough that everyone still, I think, do very well.
Jack Farley: Yeah, it's interesting that the Kimi thing, there's obviously when we talk about these, there's the demand aspect in terms of just efficiency of the models. If you don't need as much, if it's more efficient in terms of, of like KB Cash and that, that whole thing, you won't need quite as much memory. And then there's also the supply aspect of, of whether we start to see more build outs. You know, there start to be talk of availability of memory from China coming onto the market and that sort of thing. But it sounds like, yeah, regardless of that, you know, Jevons Paradox and those ideas still hold true. And regardless of these, you know, marginal change, obviously it's, it can feel especially volatile and sensitive when you have, have these, you know, memory equities that have just ran like they have like any sort of marginal shift and just with the amount of like leverage positioning, it feels very intense in the short term. But what you're saying is that regardless of that, like if you zoom out a little bit, you know, these are, these are pretty small changes on the margin.
Steve Hou: Yeah, I mean, I, I, I, I couldn't tell you, you know, in terms of sort of the stock price of these memory makers, like nobody knows, right? Everything is being priced on second or third derivatives at this point. But it doesn't seem to me at all that there's going to be any latch up on the demand for memory. And I think I fully expect to be surprised at the upside because every single time we have had a new innovation that seems to only increase essentially people's willingness to try to use these models in more, more productive ways. And every time you do that, you generate sort of longer context, you generate more data to be stored. By the way, it's not just memory, it's also storage of data. And we haven't even sort of gotten into the multimodal stuff with the voice. I think for example, over the last couple of months, one thing that people really, I think overlooked was that ChatGPT Live came up with this very, very good conversational sort of AI where you have to talk to it and you can interrupt each other. AI actually is able to interrupt you. Now you have this really powerful way of engaging with the AI where you can ask AI to go and use agents to do things on your behalf. You can interact with what you're looking at on the screen and just use voice to do it. And that also means it's a ton more data. Right. So every time you have, we haven't even touched, scratched the surface of video. Right. So I just think like, you know, every comma, like memory is ultimately a commodity. Like, I don't care, it is a super cycle, but super cycle of a commodity can, you know, also go pretty crazy. So, yeah, so, so that's how, that's sort of how I think about it. I, I'd be very, very shy about trying to make any sort of prediction about when you know, will end or how you know, but direction seems pretty clear to me for demand.
Jack Farley: Yeah, 100% cent. That being said, the last time you came on, I think at the end of the interview you're telling me that you think token efficiency is going to become really important. And it did. And so props to that. So I'm curious, how do you think the next few months play out? What do you think are the most important dynamics? It feels like one of the big ones right now is of course this discussion and deliberation around regulatory oversight of models. Especially this distinction between US frontier models versus the open way models from, from China. And you know, do we try to regulate that or even just, you know, new model launches, who gets it and when these sort of aspects. How are you thinking about all of that right now?
Steve Hou: I mean, so the way I think about it's almost a bit like, you know, the Tesla BYD situation. Right. You know, sort of the Chinese evs, ostensibly we don't see them here. Right. But they are apparently very good and they dominated the, the rest of the world market. But Tesla still here. And I expect something along that line. To happen. I think it's a little bit of, I think the reality we live in, this economic decoupling and geopolitical decoupling between the US and China that drives I think a lot of the underlying dynamic. Some of these bottleneck trades means exactly because the US doesn't want to create a strategic economic dependency on China, but buying from China. That being said, I do think the directional travel is going to be pretty similar regardless. I think there will be more efficient licensable models or cheap models that come out of the US regardless whether or not the Chinese models will be fully accessible or maybe in some sort of ways be made soft, so to speak, less accessible. You can easily imagine a scenario where the government of US put out a rule saying oh, if you few government supplies to the US federal government or work for the federal government, maybe you cannot have certain rules around safety, model safety or whatever. And then the Chinese models happen to not pass. So I think there are many, many ways in which a signal like that can be sent I think is I think pretty pretty much within expectations. That being said, I do think the efficiency gains that are made will continue to happen and China at the same time will also, now that they have caught up close to the frontier, will also be probably invest a lot more. So I wrote about this at the beginning of the year of this dual thing and US and China both doubling down on capex and state intervention and involvement probably just means that I think double the demand for overall compute and hardware. To the extent you were mentioning a Chinese memory that gets eaten up by Chinese demand by itself. So that's one thing. But then on the other hand I actually think the bigger thing to look forward to in terms of where things are headed is actually genuine enterprise adoption and the so to speak rri, right. Return on investment. I mean the consumer use case is kind of stable stakes. And the reason why by the way we see all these Chinese models being exposed, people being given out for free, so to speak, open ways models to the US is because the Chinese economy has been weak and Chinese companies historically don't have a habit of paying for SaaS, right? So you don't really have a way of monetizing the models. If you just simply create consumer use cases like a chat companion or whatever, you can't really charge that premium. But Chinese companies are increasingly willing to pay for cloud, so that's changing. And also I think in the US you're going to see I think this gradually perking up of adoption. So we already know anecdotally when you and I talk to anybody who uses AI in a small company, a small context, personal context, every single person will tell you they fundamentally change their workflow, get a lot of productivity and whatever, but you don't hear it at the aggregate and you don't see that in the aggregate economic statistics. I think what we are going to see in the coming quarters is that ironically or maybe exactly causally, because of the models have gotten so cheap, you can actually have true token maxing. Because models are so cheap, enterprises can actually let the models run wild and say hey, here's really cheap open weights model or whatever less expensive model or maybe we find some sort of smart routing thing, you can actually go experiment and do whatever you want and not worry about token budget. It's not as much anyway. And that's how actually you get the adoption of AI into the workflows. As companies discover the ways in which they can sort of create that orchestration and ownership of their own data and how the data, how their own companies talk to the AI model so that AI becomes an upgrade to their actual company's products and you create this virtual site cycle instead of everyone just being sort of siphoned dry and by the AI foundational model and sort of become the Claude whatever law firm, every company actually create a virtuous natural demand for these AI models as part of their overall business model. So I think that evolution towards that new paradigm time is what we should be looking forward to. And I think we will see more visible roi, albeit being the J shaped curve thing. It's probably going to be still slower than what people hope for, but I think signs will become more evident.
Jack Farley: Yeah, makes sense to me. Steve, appreciate you joining the show today and walking through all of that. Hopefully people learned a thing or two and I'm understand the nuances of of all this interesting data. It's a fascinating space moving fast like you're on a couple months ago and it feels like an entirely different world since then. So yeah, appreciate you coming on.
Steve Hou: Thank you so much for having me.
Jack Farley: All right, thanks. Nothing said on for guidance is a recommendation to buy or sell any investments or products. This podcast is for informational purposes only and the views expressed by anyone on the show are solely their opinions, not financial advice or necessarily the views of Blockworks. Our hosts, guests and the Blockworks team may hold positions in the company's funds or projects discussed. As always, investments in blockchain technology involve risk. Terms and conditions apply. Do your own research.