Share

By Jay Park, Kerry Wei, Denny Hua, and Erica Li

Nobody at a board meeting asks ‘should we use AI?’ anymore. Now they ask why the pilot that wowed everyone in March still isn’t in production in October, and why the serving bill tripled.

Most of the industry has spent the past three years obsessed with building AI models. But a model creates no value until it runs. Every AI application that graduates from demo to production becomes an inference workload, and those workloads are compounding fast. The bottleneck has shifted from model creation to the complex, six-month slog required to reliably serve traffic at a cost acceptable to the CFO.

AI spend will represent $1.3 trillion by 2029, growing  ~32% year-over-year from 2025-2029, driven largely by agentic AI adoption. As of 2025, 88% of enterprises had reported using AI in at least one business function, up from 55% two years prior. And the open-source model ecosystem, GLM, Moonshot, Qwen, DeepSeek, and others, now trails the closed frontier by as little as a few months, giving enterprises real optionality for the first time.

However, actually running these models at scale in production remains the single hardest problem in enterprise AI, largely because meeting the business’s demands for latency and reliability requires navigating a complex landscape of constrained GPU compute and enormous serving complexity. It is a hurdle that most teams simply are not equipped to clear on their own, especially as they often lack the rare ML systems expertise needed to solve these infrastructure challenges internally.

Fireworks: The Inference Engine for Enterprise AI

We tracked Fireworks for more than a year before our portfolio company Replit, introduced us to the team. In our first meeting, Lin made it clear they are building frontier training and inference on one platform. 

Fireworks has built a leading inference platform for enterprises serving and fine-tuning open-weight models in production. On the inference side, Fireworks provides significant cost and performance advantages, delivering roughly 10x better price-per-token economics and up to 5x faster speeds than closed-source alternatives on coding and agentic workloads. 

Importantly, its capabilities extend beyond inference to custom fine-tuning and reinforcement learning, enabling enterprises to leverage their proprietary data, workflows, context, and more, to better build, serve, and scale models tailored to their specific business needs. Its platform abstracts the complexity of more than 100,000 deployment configuration combinations into one full-stack inference platform spanning post-training, fine-tuning, reinforcement learning, and model serving, providing the tools enterprises need to scale customized models with greater speed and with differentiated economics.

Fireworks serves leading enterprise and AI-native customers, including Cursor, Perplexity, Cognition, and Genspark. Annualized revenue run-rate surpassed $1 billion, up from $280 million at its last funding round. Fireworks now processes more than 40 trillion tokens per day, having nearly tripled the daily volume of tokens served on its platform in the same period.

That’s why we’re proud to participate in Fireworks’ $1.5 billion Series D, led by Atreides Management, Index Ventures, and TCV, along with participation from existing investors including Evantic, Lightspeed Venture Partners, and NVIDIA.

Our Thesis: Inference and beyond

We always ask: how is this going to be deployed in the real world? Models will keep leapfrogging each other, but we’re prioritizing the essential plumbing that turns AI into a practical business reality. We started mapping the inference layer eighteen months ago, when we noticed our portfolio companies’ biggest line item was oftentimes serving costs.

A coding assistant lives or dies on completion latency. A real-time answer engine can’t wait on a slow model. An agent running thousands of steps unravels when the cost-per-token becomes unsustainable.

And the pattern is extending beyond AI-natives. Specialized AI is becoming the enterprise default. More companies are moving beyond one-size-fits-all AI and building their own customized models. Fireworks is helping lead that shift. 

Enterprises are using Fireworks to fine-tune open models on their private data and unique context, creating specialized intelligence for application-specific models that match or outperform generic frontier alternatives on their workloads, while running faster and at a fraction of the cost. 95% of tokens served through Fireworks today are from models that have been specialized. This is the ground on which Fireworks keeps winning enterprise deployments: performance, reliability, and economics that hold up at scale.

The Team and the Path Ahead

Lin Qiao and her six co-founders have spent their careers on exactly this problem. Qiao and several co-founders were key leaders and engineers behind PyTorch, the open-source framework underpinning much of modern AI. The broader team built and led AI / ML platforms and large-scale production infrastructure across Meta and Google. Few teams have built and operated production ML systems at that scale, and fewer still have turned that experience into a platform others can use. 

That pedigree shows up in the product. Inference is a crowded, fast-moving market, but Fireworks competes on more than just inference – its advantages compound: custom kernel optimizations that create performance advantages, GPU utilization gains that grow with volume, and a full-stack platform that goes well beyond reselling compute. This platform breadth positions Fireworks as a core layer in an enterprise’s AI stack, providing the infrastructure to power tailored solutions at scale across the business.

As AI moves from experimentation into production, Fireworks offers exactly what that shift demands: customized open models served with frontier performance, reliability, and economics. Excited to see what comes next as the company invests this new funding into scaling its platform, expanding global infrastructure, deepening its partnerships, and helping more enterprises build specialized intelligence. 

Disclaimer: For informational purposes only. Not an offer of any securities or of the investment advisory services of Prysm Capital, L.P.  References in this piece are provided solely for illustrative purposes to highlight aspects of Prysm Capital’s investment approach and do not purport to represent all investments made by Prysm Capital. Past performance is not indicative of future results.