Frontier Overhangs

Listen to this post:

There has been, over the last week, what I think is a healthy debate about the philosophy and psychology that undergirds the views of meaningful segments of the AI community, particularly those obsessed with doomsday scenarios. It is, in the end, difficult to reason with a philosophy that grants equivalent moral weight to not just all beings — human or not — who exist today, but who may ever exist in the future; this tilts the scales in such an absurd fashion towards safetyism that innovation is impossible and freedom is intolerable.

Worse, it taps into the psychology of religion, where dissent is not brooked and questioning the premise is heresy. I guess that makes me a heretic then: I reject the premise in favor of doubt in our ability to foresee the future, combined with faith in humanity figuring things out along the way. If that leads to the most fantastical doomsday scenario, then I think the State of New Hampshire said it best:

I’m referring, of course, to the question of Effective Altruism and the extent to which it is intermingled with Anthropic and CEO Dario Amodei’s insistence that We Must Pace the Frontier. I expanded on my objections to some of the philosophical implications in the last | two episodes of Sharp Tech, but Stratechery is a site about strategy and technology, and those angles deserve examination as well.

Three months ago I explained in Anthropic’s Safety Superpower how the company’s genuine belief in its safety rhetoric conveniently gave it license to pursue extraordinary goals that aligned with its business interests, including aggressive attempts to disintermediate software, collect sacrosanct customer data, and secretly sabotage would-be competitors. This analysis is in a similar vein but from the opposite direction: “Pacing the Frontier” is framed as — and I believe motivated by — a concern about safety, but it also happens to address several distinct problems faced by the frontier labs. These problems take the form of overhangs that have formed by virtue of how rapidly models are improving; slowing model improvement would reduce the overhangs.

The Capability Overhang

It was six months ago that I wrote Agents Over Bubbles, where I argued that the agentic paradigm, which kicked off with the release of Opus 4.5 in late November 2025, was so incredibly capable and so incredibly token hungry that we were not in an infrastructure bubble: we really do need all of the compute that is being built.

I do, as a matter of my job, talk about things that I do not necessarily experience personally; the canonical example is my long-running analysis of and advocacy for advertising as a business model, despite the fact I run a subscription business. I mention this because I developed this take on agentic AI before I dove headfirst into agentic coding for myself, and just as well: I have been so taken by the possibilities of making software for myself that I have started multiple new projects over the past few months, and I’ve even finished a few of them! It’s thrilling, and the possibilities seem endless. Needless to say, I believe this more than ever.

At the same time, there is one part of that Article that I’ve wavered on, and that is the importance of model company harnesses:

Specifically, I noted above that what made Opus 4.5 compelling was not the model release itself, but changes to the Claude Code harness that made it suddenly dramatically more useful. What this means is that model performance isn’t the only thing that matters: the integration between model and harness is where true agent differentiation is found.

This is a very big deal when it comes to figuring out the future structure of the AI industry and where profits will flow, because profits flow away from modular parts of the value chain — which are commoditized — and flow towards integrated parts of the value chain, which are differentiated…It follows, then, that if agents require integration between model and harness, that the companies building that integration — specifically Anthropic and OpenAI — are actually poised to be significantly more profitable than it might have seemed as recently as late last year. And, by the same token, companies who were betting on model commoditization may struggle to deliver competitive products.

One of the projects I’ve undertaken over the last month or so actually entailed — at least in one of its permutations — building my own harness, and it’s doable! It also is very difficult, at least with my level of agent-mediated capability; that noted, my go-to example in that Article about how much harness-model integration matters was Microsoft anchoring its new E7 enterprise offering around Claude Cowork, but CEO Satya Nadella broke the news in a Stratechery Interview that that was only a temporary state of affairs:

We’re using the same harness that we use in GitHub and the same thing in security, too. So we have the same harness that’s a multi-model harness in which we will rotate through — obviously MAI by default gets trained in our harness, but we will have GPT, we will have Anthropic in there and any open weight model. We will allow anyone to take any of the models they fine-tune or build. In fact, they can take an open weight model from Fireworks, tune it, put it into Copilot, no problem.

It took a while for Nadella’s claims to reflect shipping reality, but sure enough you can now choose a model for Copilot Cowork; it will be up to end users to determine how competitive Microsoft’s harness is with Claude’s, but clearly the harness and the model can be different things.

This is actually a more meaningful development than it might seem, for reasons that go back to the theory of integration and modularity put forward by the late Clayton Christensen; from The Innovator’s Solution:

The improvement of integrated versus modular systems over time according to Professor Christensen

The left side of figure 5-1 indicates that when there is a performance gap — when product functionality and reliability are not yet good enough to address the needs of customers in a given tier of the market — companies must compete by making the best possible products. In the race to do this, firms that build their products around proprietary, interdependent architectures enjoy an important competitive advantage against competitors whose product architectures are modular, because the standardization inherent in modularity takes too many degrees of design freedom away from engineers, and they cannot optimize performance…

Once their requirements for functionality and reliability have been met, customers begin to redefine what is not good enough. What becomes not good enough is that customers can’t get exactly what they want exactly when they need it, as conveniently as possible. Customers become willing to pay premium prices for improved performance along this new trajectory of innovation in speed, convenience, and customization. When this happens, we say that the basis of competition in a tier of the market has changed.

The pressure of competing along this new trajectory of improvement forces a gradual evolution in product architecture, as depicted in figure 5-1 — away from the interdependent, proprietary architectures that had the advantage in the not-good-enough era toward modular designs in the era of performance surplus. Modular architectures help companies to compete on the dimensions that matter in the lower-right portions of the disruption diagram. Companies can introduce new products faster because they can upgrade individual subsystems without having to redesign everything. Although standard interfaces invariably force compromise in system performance, firms have the slack to trade away some performance with these customers because functionality is more than good enough.

One of my go-to examples in Anthropic’s Safety Superpower was the company’s decision to predicate Fable usage on Anthropic holding onto all customer data for at least a month; this was a big deal, and I argued at the time that Anthropic was making a bet that its models were good enough to convince enterprises to give up on zero data retention:

It’s pretty significant, I think, that Anthropic is declaring that not retaining data is no longer an option, at least if you want access to their best models. Yes, today, that retention is for safety purposes only, and not for training; it’s plausible, however, that Anthropic’s lead becomes so significant that they quietly announce that they are going to train on that data as well, and companies will feel they have no choice but to go along. That additional training data, of course, will only further increase Anthropic’s lead, and all of this will be justified because Anthropic has already clearly decided they are the only ones who can be trusted to be in charge.

In fact, Fable wasn’t good enough: customers pushed back, and Fable usage stayed relatively low; when Fable 5.1 was released, the Anthropic-gets-to-keep-your-data provision was gone. This is evidence of Christensen’s theory in action: customers demonstrated the willingness to base their model-choice decision on something other than pure performance, namely, data retention policies.

This doesn’t, in and of itself, suggest that new model capabilities aren’t desired; it does, however, suggest that current model capabilities are “good enough” for customers to not do whatever is necessary to get access to the cutting edge, which reduces the value of the cutting edge to its proprietors, and gives credence to the strategy of Microsoft and others focused on separating harness and model. Pure capability no longer translates directly into a moat.

The Product Overhang

The modularization of models and harness explains why the frontier labs have what I called an economic imperative to own end user touchpoints; again from Anthropic’s Safety Superpower:

It has long been clear to me that the frontier labs have the economic imperative to move closer to the user. If you own the user touchpoint, then you have meaningful lock-in, and the best way to own the user touchpoint is to be the canvas for everything they need to do. This, by extension, means that the frontier labs are on a collision course with software companies: it’s software that owns the user touchpoint, and it’s in the frontier labs’ long-term interest to not simply be a commodity input into software but to simply replace software outright.

The key for the frontier labs, then, is to build those user touchpoints while they have superior capabilities. However, this is where Meta’s recent launch of Muse is a bearish signal. Muse is, by a significant margin, the best and most approachable personal agent product I have tried. Meta deserves a tremendous amount of credit for the product work they have put in, as well as the massive infrastructure commitment entailed in providing users with a very capable virtual machine for free. Oh, and of course they deserve credit for the Muse Spark model undergirding Muse.

That noted, Muse Spark 1.3, the most advanced Meta model, is still not state-of-the-art, and that is the bearish signal: it is good enough for a very good personal agent product, and critically, a personal agent is much stickier than a chatbot. Once you have put all of your information into a personal agent and actually incorporated it into your day-to-day life, it is much more of a challenge to change to something else. This is in contrast to Codex/ChatGPT and Claude Code: yes, you may have developed your own set of skills and understanding of how each harness works, but at the end of the day the relevant artifacts (i.e. your code) are in GitHub, and it’s not that much of a lift to point a different agent and harness at those artifacts if the alternative is better and/or cheaper (or doesn’t want to keep all of your data).

In short, model capability is good enough that compelling products — products that actually have moats — can now be built, and from a business perspective it would do Anthropic and OpenAI good to devote more of their resources to actually building such products.

The Pricing Overhang

In July I wrote Who’s Afraid of Chinese Models, where I argued that nearly everyone’s understanding of the threat Chinese models posed to the frontier labs was overstated, and an artifact of demand exceeding supply. A world with sufficient compute is one where intelligence is a commodity, and in commodity markets margin comes from a superior cost structure, which I would expect the leading model providers to have, in part because they have a meaningful lead in scaling and can apply superior AI to their infrastructure. I wrote:

All of this is to say that I think the reaction to Kimi and Chinese models generally is pretty over-blown, at least from an economic perspective. Right now there is a price umbrella that is downstream of the lack of compute; I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence.

A price umbrella is another way to say price overhang. OpenAI and Anthropic charge high prices because they can: there is so much demand for their products at their current price levels that they can barely keep up. At the same time, a huge amount of OpenAI and Anthropic’s supply does not go towards inference, but rather training, reinforcement learning, and R&D: that is compute that does not go towards reducing their price overhang. If they slowed down progress they could re-allocate computing and fully capture the market.

The Capital Overhang

I reiterated above that I don’t think that there is a bubble in terms of compute availability, but I noted in Nvidia’s Risky Business that there might be a timing problem:

This might not cost Nvidia anything in the end: if AI revenues truly take off, then the debt markets will open back up, and ultimately companies will go back to funding infrastructure investment through free cash flows. Right now, however, is the danger zone, as hyperscalers blow through the debt markets and Google at least starts to tap equity. To the extent Nvidia competes through novel funding mechanisms that, at the end of the day, draw on things like insurance floats and pension funds and other long-run liabilities that are the bread and butter of the asset managers the company is partnering with, the risk — unmarked, unlike equity — is considerably higher.

That’s why I started with 1870 and Cooke’s ill-fated agreement with Northern Pacific. Yes, the upside the deal afforded Cooke was incredible, but it was incredible for a reason: it was very risky, and pioneering new funding mechanisms only served to spread the pain when it all blew up. It’s one thing to spend all of your free cash flow; it’s another thing to tap the debt markets. And, beyond that, it’s a completely new nerve-racking thing to bring safety-seeking assets to bear. AI better deliver before it’s too late.

Anthropic is telling investors that it is profitable, but that comes with a big caveat; from the Financial Times:

Anthropic has told its backers it will be profitable this quarter, as it moves to allay investor concerns about the aggressive cash burn of frontier AI companies ahead of its blockbuster initial public offering. The company has told a small group of shareholders that its adjusted operating income will be positive for the second consecutive quarter, according to multiple people with knowledge of the matter. The measure strips out costs including stock-based compensation. Anthropic’s gross margins are above 80 per cent before accounting for revenue shared with distribution partners, including Amazon, and the cost of training its models, according to two of the people.

Excluding stock-based compensation is a pretty big caveat, but to be fair, the concern in terms of a capital overhang is cash; the bigger issue is the exclusion of training costs, which are massive (and which, in a true depiction of gross margins, would count as depreciation). That is the part that needs to be covered by revenues before the world runs out of capital, which is to say that Anthropic would be fine if only they didn’t have to pay for training!

As it stands, both Anthropic and OpenAI and the hyperscalers and neoclouds need AI revenue to increase dramatically; there is a limit to available capital, and finding that limit will end badly for everyone depending on outside capital to fund future buildouts — even if it by no means would mean the end of AI.

The Safety Overhang

I made pretty clear at the beginning that I fundamentally disagree with the framework that many in the AI safety movement operate under; that doesn’t mean I don’t think that AI safety is a serious issue, or that there aren’t already real risks. Indeed, one of the implications of there being a capability overhang is that cybersecurity risks are very real, and are going to be a massive challenge in the next few years. From Autonomy and Innovation:

The expected value for a hacker’s automated attack is always positive. If the offensive agent finds a vulnerability and creates an exploit, and that exploit fails or is itself buggy, then nothing has changed about the status quo: the exploit doesn’t work (or, perversely, makes the original vulnerability larger by virtue of its own bugs); if the agent executes the exploit perfectly, meanwhile, the attacker has gained access to the system. The attack only needs to work once for the entire endeavor to have a positive payoff.

The challenge for the defender, on the other hand, is that they need to keep the software in question working correctly, and not make the situation worse. This means that any automation has a negative expected value: successful automated vulnerability discovery and patching preserves the status quo, i.e. the software is not hacked. However, any unsuccessful patches make the situation worse, either by breaking the software or by introducing new vulnerabilities. The agent only needs to fail once for the entire endeavor to have a negative payoff.

This is the dynamic that leads to the exact situation Dalton describes, where offensive actors are fully automated while defensive systems, even if they use AI, will be incentivized to keep a human in the loop, and no human in the loop will be able to keep up with fully automated agents. Truly effective defense will mean truly trusting agents to act autonomously, but most companies won’t do that until they are forced to by regular and unremitting hacks by fully autonomous attackers.

First, it’s worth pointing out that this risk is not alignment risk, at least as that term was traditionally defined: a model being directed to do bad things is aligned; suggesting that models ought to know what is good or bad is an entirely different consideration, and I think it is very problematic that these questions have been conflated. The fact of the matter is that LLMs are, if anything, too obsequious; there simply isn’t any evidence of LLMs having a will or operating with malevolence.

Second, arguing against progress because of cyber risk was relevant before the agentic paradigm; at this point the genie is out of the bottle — and open weights models capable of attacks are already here.

Third, this reality actually makes the case for pushing the frontier, not pacing it. The capability overhang is entirely on the offensive side: it’s defenders that need models that are not only good enough to mount a defense, but to do so in an entirely automated way that doesn’t bring down the infrastructure being attacked. We’re not there yet.

In other words, when it comes to the tangible safety risk that exists today — bad actors using aligned LLMs to attack infrastructure — pacing the frontier actually increases the window in which bad things can happen.


This isn’t, of course, what Amodei and the doomers are talking about: they are worried about recursive self improvement enabling AI to improve itself, independent of human control, and I’m open to debating the risks that might result were there actually a tolerance for debate, instead of an insistence on acquiescence.

And, of course, all of the frontier labs are free to pace their own progress. That they won’t is a reminder that this is personal: Anthropic exists because Amodei and his cofounders didn’t trust Sam Altman and OpenAI; I’m not sure it’s a coincidence the demand to pace the frontier came when the frontier was, for the first time in a while, set by OpenAI.

Indeed, that’s probably the overhang that matters most of all: competition. Anthropic is fine with Anthropic being in the lead; anyone else requires government intervention. That there are safety arguments to be made that just so happen to align with their need for more time to build a moat is, I’m sure, but a sign from Silicon Valley’s newest god to its self-ordained priesthood.