Welcome back to The Cap Table Newsletter where we break down what’s actually happening in startups and private markets.
This week, I want to talk about a part of the AI stack that is becoming massively important very quickly: inference infrastructure.
For the last few years, almost the entire AI conversation centered around training. Bigger frontier models. Bigger GPU clusters. Bigger data centers. The market became obsessed with who had the most compute and who could train the largest models.
But I think the next major bottleneck in AI may look very different from the last one.
The first generation of AI products was mostly humans typing prompts into chatbots. You ask a question, wait a few seconds, get an answer back. That worked because humans themselves were the bottleneck. We can only type and consume information so quickly.
But the next generation of AI feels much more agentic. Coding agents are already starting to execute workflows continuously in the background. Voice agents are handling real-time conversations. AI systems are querying APIs, reading documents, writing code, checking their own work, and handing tasks off to other agents.
That creates a completely different compute problem.
A single agentic workflow can consume hundreds of thousands of tokens. And once agents start interacting with other agents, there’s no human bottleneck slowing consumption down anymore. Usage starts happening at machine speed.
That’s why inference infrastructure suddenly matters so much.
Recently, Two Roads Capital closed our first syndicate investment into General Compute, a company building an inference cloud designed to run AI agents, coding copilots, and real-time AI systems significantly faster and more efficiently than traditional GPU infrastructure. The SPV ended up oversubscribed in under two hours, which honestly says a lot about where investor attention is starting to move underneath the AI stack.
If you want to start angel investing, apply to our syndicate!
Follow our new and improved Instagram to hear about Trending Deals, Founder Stories, & Investor News!
Subscribe Now to get this newsletter delivered straight to your inbox every other Friday morning!
Why This Market Is Starting To Matter
Historically, most of the AI compute race was driven by training demand. Whoever had the most GPUs had the advantage. But inference introduces a very different optimization problem around latency, token speed, deployment costs, and power efficiency.
At a certain point, speed stops being a technical benchmark and starts becoming part of the product itself.
If a coding agent takes an hour to complete a workflow instead of five minutes, that changes the experience completely. If voice AI lags during a conversation, the product breaks. If autonomous systems become too expensive to run continuously at scale, the economics stop working.
That’s why the market is quietly starting to shift from: Who can train the biggest model? to: Who can actually run these systems efficiently at scale?
And those may end up being two very different winners.
You’re already starting to see signs of this happening. NVIDIA acquired Groq earlier this year in a deal reportedly valued around $20B. OpenAI also reportedly signed a massive multi-billion dollar agreement with Cerebras around inference infrastructure and launched Codex-Spark running at over 1,000 tokens per second on Cerebras hardware.
Anthropic has even shown enterprises are willing to pay significantly more for speed. Earlier this year they launched Opus Fast Mode, reportedly charging multiples more for much faster output on the same model.
That feels important.
Because for years GPUs were the obvious answer. They were flexible, available, and supported the entire AI ecosystem. But as workloads become larger, more repetitive, and increasingly agentic, specialized inference hardware starts becoming much more attractive.
Why We Backed General Compute
What caught my attention about General Compute wasn’t only the chip performance itself.
It was the infrastructure strategy.
The company is deploying specialized inference chips from SambaNova instead of relying entirely on traditional GPU architecture. Their argument is that the future agentic world requires silicon specifically designed for inference workloads rather than hardware optimized primarily around training.
But honestly, the more interesting piece to me was deployment.
One of the biggest problems in AI infrastructure right now is not just getting chips. It’s getting them online fast enough to meet demand. Most next-generation AI infrastructure projects require massive power upgrades, cooling systems, and years of construction before capacity actually comes online.
General Compute is taking a different approach.
The chips are air-cooled instead of water-cooled, which means they can fit into existing data center infrastructure without requiring massive rebuilds. The company is pursuing colocation agreements not only with traditional data center operators, but also with crypto miners looking to repurpose infrastructure as Bitcoin mining economics become more difficult.
That immediately reminded me of the early CoreWeave story.
One of the biggest lessons from the current AI cycle is that infrastructure winners often emerge from whoever can deploy capacity fastest once demand explodes. CoreWeave started as crypto mining infrastructure before becoming one of the most important GPU cloud providers in AI because they were positioned correctly when the market suddenly needed compute at scale.
General Compute feels like a very different but related bet on where the inference layer is evolving to next.
The company is also building an OpenAI-compatible API layer, which means developers can switch providers with very little integration work. That matters because infrastructure adoption usually comes down to friction. If developers can test dramatically faster inference with minimal engineering overhead, adoption can happen much faster than people expect.
The Founder Piece Matters Too
What also stood out to me was the founder background.
CEO Finn P. previously bootstrapped Fluency Academy into one of the largest online language schools in South America, scaling it to roughly $40M ARR and over 120,000 students annually before selling a stake to General Atlantic.
Before General Compute, he reportedly invested over $1M of personal capital into AI compute research before ever raising venture funding. I always pay attention when founders become deeply obsessed with infrastructure problems before the market fully realizes how important they are.
CTO Jason Goodison is a YC alum and former Microsoft engineer who worked on distributed systems for Windows while also building multiple AI startups and developer products.
I think that combination matters.
One founder deeply understands scaling businesses and operations. The other deeply understands distributed systems and developer infrastructure. A lot of the best infrastructure companies are built by teams that understand both sides of the equation.
The Bigger Picture
The inference market today honestly feels very similar to where GPU infrastructure felt a few years ago.
Still early. Still misunderstood. But increasingly critical.
The market spent the last few years obsessed with building the models themselves because training was the bottleneck. But as models move into production and agents begin operating continuously, the bottleneck shifts.
The question becomes less about who can train the biggest model and more about who can run the most useful workloads at the lowest cost and highest speed.
That’s a very different market.
And if AI agents really do become one of the largest compute workloads in the world over the next decade, the companies that lock up inference capacity early may end up controlling one of the most important layers of the entire AI stack.
My Take
The first AI infrastructure wave was about training models.
The next wave may be about powering billions of autonomous agent interactions continuously happening in the background.
That requires a completely different infrastructure layer.
General Compute is a bet that inference will not just be a feature inside existing GPU clouds. It will become its own market, with its own infrastructure, economics, and potentially its own category-defining companies.
That’s why inference infrastructure is going to become one of the most important themes to watch in AI over the next few years.
We believe General Compute has the potential to become the CoreWeave of AI inference.
👋 That’s all for now friends! See you next week.
In the meantime, Follow our instagram to see the latest founder and VC updates.
💌 Join our subscribers and sign up for this weekly Cap Table Newsletter if you haven’t already!
Also if you are interested in starting to Angel Invest you can apply to our syndicate to see our weekly deal flow!
Disclaimer: The Cap Table DOES NOT provide financial advice. All content is for informational purposes only. The Cap Table is not a registered investment, legal, or tax advisor or a broker/dealer.
