Open Weights AI Summit · Sep 01, 2026

Aaron Levie on open weights and enterprise AI

The Box co-founder and CEO explains why open-weight models strengthen the applied AI economy, how enterprises will adopt agents, and why the frontier remains radically unpredictable.

  • Co-founder & CEO, Box
  • Edited transcript
  • 39 minutes
Aaron Levie speaking at the Open Weights AI Summit
Aaron Levie · Open Weights AI Summit
Listen to the Q&AFull recording · 39:03
In his words

Highlights

All you really need is open weights to be two, three, or four months behind, and you will peel off tasks as they start to mature.

Every time you see model capability progress with open weights, you will see a directly proportional benefit to the applied AI layer.

Don't cost-optimize your automation at the start. Figure out what even works.

It might have nothing to do with where you are. It has to do with your Twitter feed.

By the time you layer in everybody's need to take their margin, you might be paying 20 times the cost of the compute's raw materials. The biggest loser is the applied AI layer. With open weights, the margins we need to worry about are Jensen's and maybe Andy Jassy's—then there's a lot of room to build innovative products on top.

Edited transcript

The conversation

Lightly edited for clarity and continuity.

Chapter 01Open weights versus the frontier

Aaron Levie (continuing)

deep down, that sort of spend—or at least profit share—goes to that ecosystem, because the alpha is the cost-saving dynamic.

With closed models, we haven't yet seen that moment. We've seen it on individual tasks. There was one example maybe two months ago in the front-end design arena. I forget which model, but it was probably a community model, and it was basically the state-of-the-art model in front-end design. That was a pretty interesting update for everybody. But it wasn't state of the art on 10 more kinds of tasks. It is possible that they benchmarked on front-end design work and eked out an extra 10 or 20 points of capability, but then maybe lost it somewhere else.

Until we see a holistically state-of-the-art model from open weights, we don't yet have the total Rubicon moment of open weights exceeding closed models. But what's interesting is that all you really need is open weights to be two, three, or four months behind, and you will peel off tasks on an ongoing basis as they start to mature. Open weights can continuously handle task categories that will shift over to them, while people will still use closed models for whatever new frontier problem they're facing.

Moderator

I think the Apple-Android analogy is really sound. There were a lot of doubts that Apple would ever be used in the enterprise, and people do now use Apple products in enterprises. Android, on a global basis, has something like 89% market share; in the U.S. it's more like 30%.

Aaron Levie

If you gave every enterprise a choice and said, “Would you like to control your intelligence or not?” 100% would say, “Yes, I'd like to control it.” The question is what happens when I add the caveat that you control your intelligence, but first you have to figure out how to run it. The ecosystem is making that a little easier with open weights.

But if I also say you're going to get something like a 10% capability discount relative to what you would get from closed models, depending on the workflow, a lot of enterprises are still going to say, “I'll just go closed.” That's going to be the benefit of closed models for a while: right now they are still ahead on certain capabilities.

Chapter 02Why open weights matter to the application layer

Moderator

What do you see in adoption? I've talked to people adopting models in enterprises, and most of them are using the cloud and closed models. When it comes to adopting open-weight models, who is happy with that setup? Is there a lot of setup and control required?

Aaron Levie

You can think about progress in open-weight capability as being directly beneficial to the applied AI layer. I put basically anything in the AI application stack at the applied AI layer. It could be customer support, legal, or anything that bridges a model's capabilities to an actual workflow in the enterprise.

There are going to be an insane number of tokens used by enterprises for, “I want to run a model. I want to own the model. I want to send in a prompt and get a response.” That will exist. The big monetizers there will, for the most part, be inference providers and maybe some frameworks around them.

But model capability in open source or open weights will most directly benefit the applied application layer. Take a pure economics perspective. Imagine an alternative world where you only had closed models, with no open models, and only two or three closed-model providers. They would have a very scarce good, and only a couple of players would have the compute or talent to produce it. It would be natural for those providers to want 80% or 90%-plus gross margins.

If you're building on top of those models, you somehow have to support the underlying model's 70%, 80%, or 90% gross margin, while also wanting the 70% to 80% gross margins that software has historically had. And obviously Jensen wants 70% or 80%, too. I may have lost track of how many layers I just added, but by the time you layer in everybody's need to take their margin, you might be paying 20 times the cost of the compute's actual raw materials.

The biggest loser in that stack would be the applied AI layer, because, unfortunately, you have the least scarce good of the three. No one is going to compete with Jensen. Maybe there is some new startup, but they'll find a way to acquire it before it gets big enough. Then there is the AI model layer, which needs hundreds of billions of dollars of capex to train the model, so you're not going to compete with them. Who gets the short end of the stick? Obviously, the software layer, because that is technically easier to build. It really hurts the software layer's ability to have long-term economics and business models.

But we appear to be entering a scenario with a very different market structure. Everyone is flirting with new business models where, at a certain scale, you might have to pay the open-weight provider some amount. But if it roughly costs what the infrastructure costs, then the margins we really need to worry about are Jensen's and maybe Andy Jassy's as well. There are still a couple of layers, but then there is a lot of room to build innovative products on top and diffuse AI into enterprises through the application layer.

It can be as simple as this: every time you see model capability progress with open weights, you will see a directly proportional benefit to the applied AI layer and to how much opportunity exists there. That is not just because deploying these models is hard, but because there are literally more dollars that can go into the application layer instead of only into raw intelligence tokens.

It is a very big deal for our industry that this middle area remains highly competitive. I have refrained from calling it “commoditized,” because it isn't yet a commodity to produce the best legal answer, the best finance answer, or the best healthcare answer. There are still only two or three providers in each of those domains. But we want lots of choice, because that creates lower prices and a little more fungibility, and it drives competition among these firms. This is all very good for us.

Chapter 03Build for traction, then optimize

Moderator

How many people here are building in applied AI? Definitely not 80%.

Aaron Levie

Okay, fine. I was being a little hyperbolic.

Moderator

For startup founders building in applied AI, would you recommend starting on open weights because that is the ecosystem to bet on long term? The technical expense is higher, but aren't tokens more efficient with open weights?

Aaron Levie

It is so hard to engineer a clean system. If you look at a specialized customer-support system, they might say, “We're probably saving 90% with open-weight models because we know the 10,000 customer-support tasks that exist, and we post-train on just those tasks.” There is no need to use a frontier model to answer, “I want to replace my laptop.” That is a small problem in your post-training.

I can't generically answer what the cost structure would be. For pure time to market, for any startup, I would always go with whatever works best for the problem—almost price be damned. Your bigger problem as a startup is getting traction, finding product-market fit, and solving problems for customers. The microeconomics of your product are a temporary issue at the moment, given how dynamic the market is.

To say it another way: if you have a startup and say you're powered by open weights, the customer literally shouldn't care. They should care only if that gives you a structural advantage that lets you underprice the alternative at the same level of accuracy. Then, of course, you should probably do it. But I would build whatever you need to get traction.

What has happened structurally—and this is very important—is that as you scale, you can peel off tasks. Harvey published a post about a week and a half ago describing a task where they processed thousands and thousands of contracts and built a table of answers about clauses or terms in those contracts. It is an insanely token-consumptive part of the product. I think they cited examples where customers spend $25,000 on one run of this approach.

At their scale, with thousands of lawyers doing that, this one feature could be a tens-of-millions-of-dollars compute problem. The feature is high-volume and frequent enough that it makes sense to post-train a model just to do that. I think they said they spent half as much on the task, with higher accuracy, by post-training an open-source model.

The collection of tasks that might have made something like Harvey structurally uninvestable suddenly changes. Again, the closed models need a 90% or 80% gross margin on inference, and then you need margin. At some point, the customer simply will not spend that much money. Open weights can completely change the cost equation.

But there was no need for Harvey to do that on day one. They could wait until they understood the surface area of their workloads and knew when to peel off those tasks. An interesting pattern is emerging: get big on whatever works from a model standpoint. Then, as you understand your use-case profile and cost profile, it starts to make sense to optimize. Each company will have to go through that journey in its own way and on its own timeline.

Moderator

That makes a lot of sense—and it is very Silicon Valley.

Aaron Levie

Yes. As with major aspects of life, assuming you can get the capital, et cetera.

Moderator

Would you give the same advice to enterprises: adopt what works, make it work well, and then optimize for cost? You've suggested enterprises are relatively cost-sensitive.

Aaron Levie

I would probably default to the same advice. It is a slightly different equation, because you don't have exactly the same time-to-market dimensions. Startups will live or die based on who gets to market first and establishes themselves. But, totally honestly, we're not going to change the law firm we use based on whether it adopted AI six months ahead of another firm. That is not going to decide whether we talk to our core outside counsel.

For the law firm, however, it will matter whether they adopt Harvey or another tool. If you're in a market where this is a differentiator, time to market and solving the capabilities that matter count for a ton. If you're simply tuning different workflows, you have more latitude in deciding among these paths.

The rough heuristic probably holds: don't cost-optimize your automation at the start. Figure out what even works. As it gets going, peel off what you can to lower-cost models. An enterprise should basically take that approach, too, though I don't have the same emotional feeling about it.

Chapter 04Where enterprise tokens will go

Moderator

You had an interesting tweet where you said you believed 99% of AI tokens would be consumed in the enterprise. What's your vision for how that happens? Reading your tweets, I imagined agents going wild inside enterprises, with people supervising them.

Aaron Levie

Minus “going wild.” I don't think of fraud detection as going wild.

We're going to wire up so many workflows to AI, even for incremental improvement. It will happen in things we've already automated, where exception handling should go to an agent instead of into a black box that nobody ever paid attention to. We'll find so many places where we can eke out another 1% improvement in fraud detection at a bank by having an agent parse something.

For example—and maybe I am on a very low-grade mode at my bank—I get alerts saying, “Are you really taking that Uber?” If you looked at my credit-card statement, I do three things: I use Uber, I use DoorDash, and maybe one other thing. Yes, that is really me in the Uber. For the 9,000th time, I am taking this Uber.

They clearly have a deterministic system with some rule that says, “This is weird because you typed in a new address,” or the amount is slightly different. At some point, somebody at the bank will say, “Our customers hate this experience. We'll wire up an agent to review these things and cross-correlate them with some other data points that require a little more probabilistic understanding.” They won't want to figure out a bespoke machine-learning model for it.

We're going to throw unbelievable numbers of agents at all of these things: bank transactions; reading every document; looking at every log and every security incident; looking through all of our code 10 times to see whether there are security flaws. That is where most of the tokens will go in the world. We are very early in wiring up those systems.

Chapter 05Adoption across industries—and its limits

Moderator

There must also be a lot of inertia in enterprises. What does the adoption distribution look like? Which enterprises are at the forefront of AI adoption, and which are behind? I used to work with restaurant chains; I can't imagine they are at the forefront.

Aaron Levie

But that is where the applied layer makes sense. Close to home for you, I think Toast packaged a number of AI products. It won't show up as a restaurant “adopting AI.” They'll say, “I'm going to adopt my new phone-call system.”

I've probably called five restaurants in the Bay Area in the past six months that are all using the same kind of front-end voice bot. They probably have more AI in that workflow than we do in our call center at the moment. It will show up in interesting ways. My hit rate for predicting an AI adopter is very low based on the normal heuristics.

With cloud adoption, it was a little easier: regulated or not regulated was the simplest way to predict it. But some banks are among the biggest adopters of AI because they have to undertake massive projects: upgrading banking systems, migrating legacy platforms, and handling customer-support interactions.

Almost all industries are moving this way. Excluding tech, which will be ahead of everybody else, information-heavy businesses may be a little further along—legal, law firms, banks—because the upside of AI is so high. Manufacturing may be a little less so, comparatively, because that is more traditional automation. But it is going to roll through everywhere.

Moderator

Information-heavy businesses are where you want to start. You also wrote about the interruptions in a business process—when you're waiting for feedback from a customer, for example, and AI naturally can't do anything. The more interruptions you have, the harder it can be to apply AI.

Aaron Levie

Yes. That is why I think some of the diffusion—or the takeoff scenario—tends to be overestimated. We can still only develop drugs so quickly. We can still only put up a building so quickly. We can still only borrow capital so quickly. Those are the bigger gaps.

Smaller gaps look like a sales rep waiting for a customer to respond. Anyone who has tried to automate as much of a sales motion as possible knows the parts you tend to automate: whom to contact and when, the pre-work on a customer and better insights, and building sales collateral. You can automate all of that. But at the end of the day, I can only get a customer on a call at the pace at which they want to get on a call.

All the agents in the world—other than perhaps reaching the customer at the right moment with the right message to increase urgency—don't change the fact that the workflow is paused for a week. I can't automate the sales motion. I can automate the information-heavy parts of it, but the human element remains. That is why a lot of this work looks different from coding: there is much more back-and-forth interaction in these workflows.

Chapter 06Coding agents and the diffusion of new tools

Moderator

Coding moves super fast. What kind of adoption have you seen for coding tools in enterprises? In Silicon Valley, it feels like almost full adoption.

Aaron Levie

Definitely full adoption in the Valley. But even there, there are unlimited tiers of maturity. Some leading companies are still working very differently from the rest, along with very small startups doing this out of sheer financial necessity.

Most people finished rolling out Cursor about a year ago. Then Claude Code happened, and now they're doing that. Then it's, “No, actually, the new thing is to have agents in Slack that you send tasks to.” And you're thinking, “Oh man, I might have to adopt that pattern.” We're always evolving.

Now imagine the version outside Silicon Valley. Last week, they adopted Cursor, and then they found out that Cursor won't support OpenAI, and now they're saying, “Oh, shit.” There will be cohorts rolling out at different stages for each of these things.

Moderator

How far outside Silicon Valley is adoption—about a year behind?

Aaron Levie

It is hard to say. Citi and Goldman Sachs did press releases with Cognition a year ago, so I can't claim that outside Silicon Valley they're always a year behind. You can get one person in one of these big companies who happens to be incredibly tapped in, and they'll roll out the most advanced things.

The world is connected enough that I can't quite frame it geographically. It might have nothing to do with where you are. It has to do with your Twitter feed. It might literally come down to: are you reading the right tweets? That causes your company to accelerate—or not.

Chapter 08Agents as security principals

Moderator

A former Box engineer asks: “I've seen how important security, permissions, and enterprise content are at scale. As AI agents become primary users of enterprise software, what needs to fundamentally change in infrastructure, APIs, identity, and access control?”

Aaron Levie

The past couple of months of sandbox-escape games have been pretty interesting. They show that these agents are heat-seeking missiles for every tiny crevice they can sneak into and roam through. They will find every crevice they can.

We need to work on alignment as much as we possibly can, but the only answer is to make sure there is literally no way in. For as much talk as there is about the OpenAI side of an escape, the other side is: what was the targeted company doing from a security standpoint? It should literally be impossible to break into another system.

The problem is that most systems can be broken into with enough compute, phishing, or other exploits. We need the models themselves to get better and better from a safety standpoint, but there will be upper limits. There will be bad actors who can jailbreak a model and turn it into something they want to use. You have to defend against that with great encryption, great security, well-maintained access controls, and very locked-down systems. Those are the kinds of things enterprises will have to invest in.

Chapter 09Can closed frontier labs stay ahead?

Audience member

Why do you expect frontier models to remain differentiated enough, for long enough, to matter relative to the aggregate of open weights?

Aaron Levie

I am mostly, honestly, extrapolating from the current situation. I haven't done a deep analysis. If anything, my view is probably the consensus view.

I would love to have the contrarian view, because it would be more fun to argue the other side. But it appears that whoever can pay for the top researchers, the best data, and the most compute can sustainably stay ahead. We haven't yet seen a structure that replicates that in an open-source capacity.

I don't have a good sense of how much distillation is or isn't happening. But let's say distillation is even remotely a part of what is happening. By definition, that means you're already a derivative of the frontier model, so you're always going to be downstream to some extent.

Audience member

But conceivably you're a break-even or unit-cost-positive derivative. What do you believe about the unit economics of the frontier labs? I think they're all in negative gross-margin territory.

Aaron Levie

I think Anthropic is very much not negative gross margin on inference.

Audience member

On inference, sure.

Aaron Levie

Then all you have to believe is that inference becomes big enough to overtake your training costs, and these look like pretty good businesses—at least for one or two players. I don't think I would like being the seventh closed lab.

Audience member

Sorry to dominate, but you have to train relatively infrequently and make margin on those tokens for quite some time, without being distilled, for that to work. In the scenario where enough distillation happens quickly enough, the open-weight players kill the golden goose. There is so much application software to write that I don't actually care; we can spend years on static models and be fine, per your margin comment earlier.

Aaron Levie

First, the scenario you're painting is certainly why distillation has become such a heated topic, and why we've seen more attempts to close it down. Clearly, you want to amortize your last training run for as long as possible, because that is where you can make as much as possible.

Some people debate how important distillation is to all of this. Some people think China will simply catch up in general. There are other views, too. Maybe I'll change my mind.

Are you going to short Anthropic at its IPO?

Audience member

I used to be an IPO banker. I don't touch that crap, so I have no expertise.

Aaron Levie

Okay.

Audience member

I bought Amazon at three in 2003, because I'm the old sage. I expect to buy Anthropic at three in 2032.

Aaron Levie

At three, when it's $20,000? Would that be bullish at that point?

Audience member

Absolutely.

Aaron Levie

I see. Honestly, I think every possible scenario is in play right now. This is the least predictable industry in the history of tech—other than that, if you have compute, you're set. No matter what, Amazon is going to make money. TPUs will make money. Jensen makes money. We just don't know what happens between us and them.

Chapter 10Small models and businesses around open weights

Audience member — RJ, venture investor

We're a venture fund that only invests in companies with an open-source product. Comparing international and domestic models, we've seen much more usage from models coming from outside the U.S. But with small language models, we're seeing Apple and other U.S. companies dominate applications on mobile devices. Where are you seeing other commercial applications, and what are you bullish about?

Aaron Levie

Small models are not something I'm as close to, because most of my focus is on enterprise workloads. The obvious applications are on-device consumer experiences. I don't have a good sense of enterprise workloads beyond maybe IoT-related uses, typing, robotics, and so on.

The only way you are really going to play in small models is at the applied layer: what is the ultimate device or application you're selling?

Around bigger open-weight models, multiple types of businesses will be built. The most obvious category right now is post-training infrastructure and reinforcement learning, which appears to be producing very compelling businesses.

Some people strongly believe every company in the future will have its own model. Some believe every team will have its own model. You can take this to very extreme conclusions. It is impossible to know exactly how it plays out, but there are compelling businesses around making these models ready for the enterprise and for particular domains. Then there will probably be businesses in tooling, inference, and so on.

Chapter 11Choosing model size—and vertical versus horizontal AI

Audience member — Tarquin, ActionMeet

We work with many enterprises on day-to-day business operations, and we haven't seen them use top-of-the-line models. They use open-weight models or smaller Gemini or Microsoft models to run day-to-day operations. We use larger, more recent models to take standard operating procedures and convert them into workflows, but we don't use those models to run the workflows. What are you observing, and how do you think about using them?

Aaron Levie

The only reason you need a larger model is probably some degree of very complex orchestration, or pulling in and pushing out adjacent tasks. We're only six to 12 months into understanding the shape of that in a process.

If you're answering customer-support requests, you're totally fine with the lower end of intelligence. I'm not trying to disparage those models. A smaller Gemini model is going to be perfect: the customer-support request comes in, and—boom—here's the answer. It is fast, cheap, efficient, and highly capable.

You start to need the next level when the agent has to pull up a computer, do some code on the fly, write a skill for processing data, or handle some edge case. Most workflow systems in the enterprise don't even support that. They don't have the ability to run a container or provide an environment in which the agent can operate. We're very early.

Your assessment is probably very consistent with what is actually happening in the enterprise. Other than actual engineers and chat systems for team work, you're probably not seeing frontier models in a real business process at the moment.

Audience member — Tarquin, ActionMeet

Do you think vertical AI companies will rule the enterprise, or horizontal AI infrastructure?

Aaron Levie

It is a total coin toss. You can look at one group of customers who say, “I want something purpose-built because I don't want to have to think.” Then you look at other companies that say, “I already have a team of 20 engineers. I need to give them something to do.” You can't really predict it.

SaaS had exactly the same dynamic. There have been incredible horizontal software outcomes and incredible vertical software outcomes. AI will produce roughly the same mix of scenarios...

[The recording ends mid-answer.]