The Vertical Intelligence Company
Ironically, the best model is becoming the wrong thing to build a company around.
That sounds strange at a moment when model capability is rising so fast. Every few weeks, a new system appears that is cheaper, faster, more open, or better at a task that looked difficult six months ago.
Founders naturally want to attach themselves to the winner.
But the winner keeps changing.
Think about the stack that should matter to every entrepreneur building with AI: an open-source model trained on domain data, frontier models sitting behind a router, and a system that sends each piece of work to the intelligence best suited to handle it. The combination can produce the same or better outcomes at a lower cost than sending every request to the frontier.
This is not just a cheaper way to build an AI product.
It changes where enterprise value lives.
If the underlying models continue to improve and their prices continue to fall, raw intelligence becomes less defensible. Access spreads. Performance gaps narrow. Today’s advantage turns into tomorrow’s API option.
The durable company will not be the one that rents the smartest model first. It will be the one that builds the best system for turning changing models into a better customer outcome.
The model is becoming a component. The learning loop is becoming the company.
The Stack Has Split
The first generation of AI products was built around a simple idea: choose the strongest available model you could afford, place an interface in front of it, add some context, and sell the result.
That was enough when capability was scarce and the model did most of the visible work. It is no longer enough now that intelligence has become a portfolio.
The emerging stack has at least three layers.
At the base is an open-weight worker. It handles the high-volume work: reading, drafting, extracting, classifying, searching, reconciling, and calling tools. It can be hosted with more control, adapted to a domain, and run at a fraction of frontier cost.
Above it sits frontier intelligence. The frontier is not removed. It is used selectively for the hardest reasoning, the ambiguous exception, the final review, or the moment when an error would be expensive.
Between them sits the routing and evaluation layer. It decides which model should perform which task, measures the result, manages context, preserves state, calls tools, escalates uncertainty, and records what happened.
Open models do the volume.
Frontier models supply scarce judgment.
The system decides where each earns its keep.
Fireworks and Harvey recently published a useful demonstration of this architecture on a 100-task slice of Harvey’s Legal Agent Benchmark. Their open-model worker called a frontier model only when it needed advice. According to their reported results, the hybrid system passed 18 tasks completely at a total inference cost of $368, compared with 14 tasks and $954 when the frontier model ran end to end.
A separately post-trained open model also improved its all-pass result at roughly unchanged inference cost. The sample was limited and the work was published by the vendors involved, so it should not be mistaken for a universal law. But the mechanism is the important part: open weights at the core, frontier intelligence called only where it changes the answer.
This turns model selection from a product commitment into an operating decision.
That is a much stronger place to build.
Intelligence Becomes a Variable Input
Software companies once chose a database, a cloud, and a programming language.
Those choices mattered, but customers did not buy the stack. They bought the capability the whole system delivered.
AI is moving toward the same structure, only faster.
A model is an input into production. Different work deserves different intelligence. A routine contract extraction does not need the same reasoning budget as a novel regulatory question. A first-pass support response does not need the same model as a threatened enterprise renewal. A common code migration does not need the same model as a security-sensitive architecture change.
Sending everything to the frontier is like staffing every task with the most expensive expert in the firm. It may feel safe. It does not scale.
The whole reason we built companies in the first place was in recognition of this fact: work can be orchestrated, and there is profit in managing work intelligently.
The router creates a new form of managerial leverage. It can assign cheap cognition to common work, specialized cognition to repeated domain problems, and frontier cognition to the narrow band of tasks where it changes the outcome. Over time, it can learn the boundary.
That boundary is valuable.
Factory, an agent-native software development company, reported that routing work across open and frontier models lowered its average task cost by 30% to 40% in private preview. Its data suggested that roughly a third of tasks needed the frontier, a third could use the cheapest reliable option, and a third belonged somewhere between. The specific percentages will vary by domain, but the economic lesson travels: the cost of intelligence should match the difficulty and consequence of the work.
This is how AI moves from a premium feature to a labor layer.
Companies are beginning to treat tokens as a substitute for labor. The important question is no longer “How do we get employees to use AI?” The important question is “How do we turn tokens into labor at the lowest sustainable cost?”
Tokenomics is the future.
When every task requires the most expensive model, the product remains constrained by inference economics. When the system can blend models, improve the cheaper worker, and reserve the frontier for exceptions, the company can automate more work while protecting margin.
The addressable market expands because the cost of serving it falls.
The Moat Moves Into the Loop
Falling model costs are good for builders.
They are also dangerous for weak businesses.
If a product is only a thin interface around one model, every model release threatens it. The vendor may add the feature. A competitor may switch to a better API. The customer may build the workflow internally. Capability rises, but the company captures little of the increase.
The answer is not to predict which model wins.
The answer is to own the assets that become more useful no matter which model wins.
I see five assets that the companies I am building and investing in can compound.



