header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Fireworks, spun out from Meta, discusses Open Source vs. Closed Source: Who Will Prevail?

Read this article in 44 Minutes
Open Source Catching Up with Closed Source, the Real Competition Has Just Begun
Video Title: Silicon Valley Insights x Fireworks Co-founder Benny Chen: Open-Source Models, Token Growth, Inference Optimization, and Model Customization
Video Author: Silicon Valley Vector
Editor: Peggy, BlockBeats


Editor's Note: Against the backdrop of open-source model acceleration approaching proprietary state-of-the-art models and continuous inference cost reduction, industry discussions are shifting from "who has the most powerful model" to "who can deploy the model into production at a lower cost." However, as model capabilities converge and Token consumption growth becomes a consensus, a more fundamental question arises: what enterprises are truly willing to pay for, cheaper model inference or dedicated intelligence that can reliably accomplish specific tasks?


Recently, host Cao Qingyun from "Silicon Valley Insights" had a dialogue with Fireworks AI co-founder Chen Yufei. Positioned between models and enterprise applications, Fireworks primarily provides clients with open-source model inference, performance optimization, and custom services. Rather than merely discussing whether open source can catch up with proprietary models, Chen Yufei's observations are more closely aligned with real-world workloads: where Tokens flow, why enterprises pay, and what is still lacking as models transition from concept validation to production.



In this dialogue, Chen Yufei breaks down the question of "who will win between open source and proprietary" into a set of more fundamental structural issues: whether Token growth can translate into revenue, if general capabilities can replace vertical accumulation, whether low-cost models can pass enterprise evaluations, and how inference platforms can capture value between cloud providers and application companies.


First, the scale of open-source model usage and commercial value is diverging. In the past, the ability to catch up and inference prices were the main indicators of open-source competitiveness; today, the Fireworks platform processes approximately 40 trillion to 50 trillion Tokens daily, and the actual usage of open-source models has rapidly expanded. However, free traffic, promotional subsidies, and model price differences may overstate Token statistics. Clients may heavily invoke low-cost models but still allocate their highest budgets to the best-performing proprietary models. This indicates that the next phase of open source is no longer just about expanding traffic but proving that it can achieve or even surpass cutting-edge models in high-value tasks and translate cost advantages into willingness to pay.


Second, general models and vertical models are starting to evolve in different directions. In the past, each upgrade of cutting-edge models could directly phase out a batch of fine-tuned models; now, vertical applications in law, medicine, programming, etc., are accumulating more detailed evaluations, data, and workflows, with their optimization goals gradually diverging from the cutting-edge labs. General models need to raise the ceiling on capabilities, while vertical models need to reliably deliver results within limited scenarios. The former can solve more widespread problems, while the latter better understands how users define "correct." This implies that the barrier for vertical companies is not just having a customized model but continuously translating industry demands into evaluation systems and iteratively migrating with base model updates.


Third, the bottleneck for enterprise AI implementation is shifting from model supply to evaluation capability. In the past, enterprise concept validation often relied on trial experience and subjective judgment; now, as AI enters production processes such as call centers, legal research, and medical assistance, relying solely on "looks good" is no longer sufficient to support procurement decisions. Enterprises must know on which tasks the model is effective, when it fails, and how much cost and quality will change when switching from closed source to open source. Evaluation is no longer an auxiliary tool, but the infrastructure that connects procurement, training, and production deployment. Those who can define tasks, establish test distributions, and continuously update standards are the ones who truly control the model selection.


Fourth, the value of the inference platform is shifting from "selling cheap computing power" to organizing models, hardware, and workflows. In the past, inference optimization was mainly understood as reducing the cost of a single token; now, caching, task partitioning, model routing, and context management can directly impact task completion rates. Different models no longer have to compete for the same position but can serve as executor and advisor separately. This is also the basis of Fireworks' business logic: not building heavy asset hardware internally, but binding revenue to actual customer model usage through training, customization, and continuous inference. However, the main competitor on this path is not a single new cloud company but large cloud providers that can simultaneously control computing power, software, and customer access.


Fifth, the rise of open-source models may not necessarily weaken infrastructure demand but may instead compress the model tier premium and further push value towards inference and computing power. Tech giants continue to increase capital expenditures, not just looking for short-term returns on investment but assessing long-term risks of missing the AI cycle. However, for new cloud companies that rely on external financing, debt costs, project delays, and return cycles will still pose more direct constraints. The fact that models are becoming cheaper does not mean that building and operating AI systems will also become lighter in sync.


If this conversation is condensed into one judgment, it is this: the convergence of open-source models is just the beginning, and the next stage of AI commercialization will revolve around evaluation, customization, inference efficiency, and customer workflows.


The original content is as follows (slightly restructured for readability):


TL;DR


· Open-source models are rapidly catching up with closed-source models, but Token growth does not equal revenue growth, and enterprise budgets still prioritize the best-performing models.

· Distillation can only help open-source models narrow the gap in the short term; long-term competitiveness still depends on independent training, evaluation data, talent, and computing power.

· General model upgrades no longer inevitably replace vertical models; barriers in scenarios such as law and medicine are shifting from model capability to evaluation, data, and workflows.

· The key for enterprise AI to move from PoC to production is not to increase model choices but to establish an evaluation system that can measure quality, cost, and failure boundaries.

·The optimization of reasoning has expanded from reducing the cost per token to caching, task splitting, and model routing, where the system design itself can enable the combination of multiple models to surpass a single leading-edge model.

·Custom models are more suitable for vertical SaaS platforms covering a large number of similar customers, as a single enterprise typically lacks a sufficiently broad data distribution and continuous evaluation capability.

·Fireworks' competitiveness lies not in owning GPUs, but in connecting training, customization, and continuous reasoning. However, the real long-term rival is still the cloud giant that can cover the entire tech stack.

·The popularization of open-source models may not necessarily weaken hardware demand; instead, it is more likely to compress the model-level premium and redistribute industry value to computing power, inference infrastructure, and enterprise workflows.


Key Points of the Interview


Open Source Will Continue to Catch Up, But Distillation Is Not the Long-Term Answer


The Fireworks team mostly comes from Meta's PyTorch ecosystem. Based on their experience with the development of past software such as operating systems and databases, the team has long believed that competitive open-source solutions will eventually emerge for software of significant value. Large models may follow a similar path; it's just that this round of catching up is happening faster than Chen Yufei originally anticipated.


On one hand, the annual recurring revenue of closed-source model companies like OpenAI and Anthropic is still growing rapidly; on the other hand, open-source models from China and the U.S. are rapidly narrowing the capability gap and driving down inference prices. The relationship between open source and closed source is not a zero-sum game of one side growing while the other inevitably declines.


Regarding distilling some open-source models through closed-source model outputs, Chen Yufei believes that this approach can help models achieve a high level in the short term but is unlikely to become a reliable long-term path. Closed-source vendors can gradually patch APIs and security mechanisms to reduce the possibility of inference processes and related data extraction.


However, the progress of open-source models does not solely depend on distillation. As long as the evaluation system continues to improve, data costs keep decreasing, and there is enough talent, GPU, and engineering capabilities, open-source teams can still independently train competitive models. Companies like Meta with computing power and talent reserves have no reason to maintain a significant gap with closed-source leading-edge models in the long run.


This doesn't mean closed-source models will lose the market. Chen Yufei compares the two to Apple and Android: the open ecosystem may capture a larger user base, while continuously leading closed-source products can rely on stable user experience, brand recognition, and enterprise trust to retain high-value customers with lower price sensitivity.


Enterprise procurement often tends to be conservative. The decision-making psychology described as "nobody ever gets fired for buying IBM" exemplifies this. As long as customers still believe that purchasing top-tier closed-source models is more secure, companies like OpenAI and Anthropic can maintain a certain advantage. User mentality and market entry capabilities may be harder to change compared to short-term technology gaps.


Chen Yufei expects that the ARR of closed-source model companies will continue to grow, but this number may not fully reflect the ultimate economic benefits they receive. For example, some revenue may need to be shared with cloud providers, and an increase in ARR does not necessarily translate into the same proportion of accounting revenue and profit. Whether the bargaining power of closed-source models has decreased remains to be seen until relevant companies disclose more complete financial information.


In the long run, open source may achieve a larger adoption scale, while closed source continues to capture demand with higher profitability. What determines the business boundary between the two is not only the model's capabilities but also factors such as brand, channels, customer trust, and revenue-sharing methods.


Token Flowing Towards Open Source, While Revenue Still Follows Effect


According to Chen Yufei, the Fireworks platform currently processes approximately 40 trillion to 50 trillion Tokens per day. Based on publicly available information, this scale has surpassed the enterprise API traffic disclosed by Gemini and OpenAI.


The main demand on the platform comes from programming, healthcare, collaborative work, and deep research. Among them, more and more tasks that were not originally considered programming are being redefined as "programming problems."


Typical examples include PowerPoint, Excel, and other office software. When models can manipulate these software through code or structured tools, existing code generation, tool invocation, and reinforcement learning methods can be transferred to office scenarios. Many vertical SaaS companies are also encapsulating industry tools into environments that models can call, and then through reinforcement learning, enabling the models to become familiar with specific workflows.


Chen Yufei roughly categorizes the current demand into two types: programming and deep research. Different vertical applications will combine the two to create dedicated Agents for scenarios such as law, healthcare, and office.


He is particularly optimistic about collaborative office products for non-technical users. Previously, most industry resources were invested in "STEM demand" such as mathematics and programming, but the demand for tasks oriented towards a wide range of knowledge workers, such as creating presentations, processing documents, and making videos, may grow faster.


The development of Agents that operate computers has progressed slower than previously expected. Over a year ago, such products were believed to be able to bypass complex APIs and interact with computers directly like humans; however, in practice, in order to reduce costs, more workflows have been shifted to plain text. Text models are cheaper and easier to integrate into existing processes. With the decreasing cost of virtual machines and improvements in operating environments, there may still be more room for computer operation Agents, but the specific optimization path is still being explored.


However, Token growth does not directly represent business value. Public model routing lists are often influenced by free quotas and promotional activities, and the same number of Tokens may correspond to completely different prices. The traffic generated by low-priced models does not hold the same significance at the revenue level as the traffic generated by expensive cutting-edge models.


Chen Yufei believes that revenue is a better reflection of customer choices than the number of tokens. When a business is insensitive to price, budgets will still flow to the model with the best performance. Therefore, when Fireworks assists clients in customizing models, the primary goal is usually not to achieve a "good value for the price," but to meet or even exceed the performance of top proprietary models for specific tasks.


A similar divergence between usage and revenue exists in the software market. Token consumption is likely to continue growing, but if commonly used models keep decreasing in price, lower prices may stimulate more usage without necessarily expanding market revenue in sync. Just as the widespread adoption of electricity does not mean all profits go to power generation companies, the growth in AI usage cannot directly answer which layer of the industry ultimately captures the value.


In the near future, the usage of open-source models may continue to rise, but whether related spending will expand in parallel depends on whether businesses are willing to invest resources in customization. If more companies can challenge proprietary models with customized open-source models for specific tasks, traffic and revenue will further shift towards open source; however, if customization costs, failure rates, and organizational capabilities continue to pose obstacles, proprietary models will still handle most high-value demands.


General Models Raise the Ceiling of Capability, Vertical Models Accumulate Task Specificity


Every time a new cutting-edge model is released, the market raises a question: Will models that enterprises invest data and funds to fine-tune quickly lose their value?


Chen Yufei stated that indeed, by 2025, many custom models have been replaced by the next generation of base models, but by 2026, this situation has significantly diminished. In his view, this is a positive signal of the gradual stabilization of the commercial value of vertical models.


The reason is that general models and vertical applications are optimizing in different directions. Frontline labs need to demonstrate that models can solve more challenging mathematical, scientific, and drug development problems; vertical companies in law, medicine, and other industries focus on specific user needs, task completion rates, citation accuracy, and workflow reliability.


"Doing a good job of an agent in a vertical field and solving the Riemann Hypothesis may be two completely different tasks."


Take the legal AI company Harvey as an example; its advantage comes not only from the base models used but also from the decomposition of legal tasks, evaluation systems, data processing, and product workflows. The same applies to medical scenarios, where different products need to establish specialized standards around how doctors work, data sources, and accuracy requirements.


These evaluations and data do not automatically become obsolete with the release of new models. After upgrading base models, vertical companies can migrate their existing training environments to the new models without the need to re-understand the entire industry. As evaluations become more detailed and data cleansing processes mature, vertical models form a continuous accumulation of capabilities around specific tasks.


In theory, leading model companies can also concentrate resources into the legal or medical markets and beat existing products in a single area. However, this approach may not necessarily align with their business goals. General model companies need to serve a broad client base, and if they invest a large amount of resources in a vertical sector, they may struggle to support the market scale and valuation they are pursuing.


Yufei Chen compares general models to a chain restaurant that caters to a broad range of tastes, while vertical models are more like restaurants focusing on a few specific dishes. The latter does not need to solve all problems; it only needs to outperform general solutions significantly in the target task to potentially establish an independent premium space.


This type of imbalanced yet highly specialized capability is also referred to by him as "Jagged Intelligence." A legal model may not be able to solve the Riemann Hypothesis, but it can outperform larger-scale general models in a certain class of legal tasks. For enterprise customers, the ability to reliably complete tasks is often more important than whether the model has a broader range of capabilities.


Not all companies are suitable for custom models. Yufei Chen believes that the most suitable clients are vertical SaaS companies that serve a large number of similar institutions, rather than individual hospitals, law firms, or end-user enterprises.


Vertical SaaS platforms can reach a large number of customers, understand common needs among different institutions, and establish an evaluation system covering more scenarios. Single enterprises often lack a broad enough data distribution and continuous evaluation capability. Instead of independently training models, they should solidify internal processes into skills, tools, or agents and then call external base models to complete tasks.


In custom reinforcement learning projects, the GPU is a major cost, and data and training environments are equally critical. The team needs to check whether the model is engaging in reward hacking, meaning exploiting evaluation vulnerabilities to achieve high scores without actually completing tasks. They also need to ensure training environment stability and scalability.


The initial environment does not necessarily need to reach hundreds of thousands of instances. Yufei Chen suggests that starting with thousands of training environments may be enough to kickstart reinforcement learning because the model will repeatedly explore and generate a large number of trajectories during training. Assuming each training step runs 128 times on one environment, a thousand initial data points could generate approximately 128,000 training trajectories.


After a custom model goes live, it still needs to "revisit." When a new open-source base model is released, the platform can migrate the existing training and evaluation system to the new model. As long as the evaluation criteria and task environment are fixed, the migration itself may not be time-consuming. What truly generates ongoing work is customers adopting new scenarios, workflows, and more challenging tasks.


Therefore, the core asset of a vertical company is not the weights of a single-generation model but the data, evaluation, and workflows that can be repeatedly migrated along with base model updates.


Enterprise AI Stuck in PoC, Lacking Evaluation Rather than Models


When choosing between open source and closed source models, enterprises need to prioritize trust over just performance and price.


Since enterprise contracts usually last one to two years, buyers need to assess whether the vendor and its ecosystem can support long-term business needs. While open source models may be cheaper and more flexible, closed-source vendors still have an advantage in brand reputation, success stories, and market awareness.


Chen Yufei believes that the most obvious shortcoming of open source models and their vendors currently is not technical but rather the ability to enter the market. Closed-source model companies can continuously strengthen user awareness through new capability demonstrations and research results, while the open source ecosystem lacks a unified marketing and customer communication system.


Deployment methods are also evolving. Early enterprises focused more on on-premises deployment, but now more workloads are moving to the cloud. As long as vendors can establish trust in security, virtual private clouds, and other areas, multi-tenant cloud services are usually more cost-effective than enterprise-owned GPUs.


A single enterprise may not be able to keep a batch of GPUs highly utilized at all times, but a platform can aggregate the needs of different customers, improving hardware utilization efficiency. As the proportion of inference costs in enterprise operational expenses rises, the cost advantages of shared infrastructure become more apparent.


As enterprises move from proof of concept to production, one of the biggest obstacles is the lack of rigorous evaluation. Even though many companies have already used AI in crucial areas such as call centers, the way they assess model performance is still largely based on trial and error, rather than establishing repeatable testing standards.


Without proper evaluation, enterprises cannot reliably compare open source and closed source models or determine whether model updates have truly enhanced the business. Therefore, Chen Yufei believes that talent capable of designing AI evaluations is still scarce. A reliable evaluation system could help enterprises complete model replacements and save long-term expenses far exceeding their development costs.


This is no different from unit testing in traditional software companies. In the past, SaaS companies needed to run a set of tests before software delivery; today, vertical AI companies need to establish a set of evaluations to ensure that the agent meets the required standards before deployment. The focus of industry work is shifting from merely "writing tests" to "writing evaluations," but the core remains transforming product quality into a measurable, accumulative organizational capability.


The development of third-party data companies can also serve as an indicator of the adoption of enterprise AI. If their customers continue to concentrate on a few cutting-edge labs, it indicates that the space for ordinary enterprises to establish independent AI capabilities is still limited; if revenue sources gradually diversify, it may mean that more companies are starting to purchase data, set up evaluations, and train their own models.


What enterprises truly lack is not more models, but a set of standards that can define the correct ones, identify failures, and support procurement decisions.


From Inference Optimization to System Scheduling: Reducing Token Costs


Fireworks' first core business is inference optimization. Chen Yufei believes there is still significant room for improvement in this area, as the efficiency of a model right after its release usually significantly differs from its performance after long-term optimization.


The problem the platform aims to solve is how to shorten this optimization cycle so that a new model can be as close to its optimal state as possible when launched. Apart from the model runtime itself, caching, scheduling, and supporting infrastructure all affect the final cost. Much of the optimization comes from a large amount of detailed engineering work, making inference optimization a labor-intensive business.


Compared to the first-party APIs provided by model developers, Fireworks simultaneously serves a large number of basic models and custom models, enabling the observation of various traffic patterns and workloads. Model vendors need to first stably serve their models to a broad user base, while third-party platforms can perform more detailed optimizations based on the characteristics of specific customer requests.


Model routing is not only aimed at cost reduction. In a joint experiment between Fireworks and Harvey, the system assigned one model as the task executor and another model as the consultant, breaking down long-context tasks into several shorter subtasks. Since the effectiveness of large models may decrease as the context lengthens, proper division of labor may result in the system's overall performance exceeding that of using the strongest model alone.


Chen Yufei likened this process to CPU performance optimization. Developers need to understand the processor's characteristics and then divide the program into workloads suitable for that processor. Similarly, in AI systems, basic models can be seen as a type of text processor: the team first needs to understand the capabilities of different models, then design the task framework, and finally place the appropriate models in the right positions.


Hardware optimization cannot focus solely on the individual metrics of GPUs, HBMs, CPUs, and networks. At the beginning of training, a model is usually designed considering the computation and memory ratio of the existing hardware. Therefore, migrating a model from an NVIDIA GPU to a similar AMD GPU is relatively easy, while migrating to a different architecture ASIC may require a significant amount of work.


Whether a new hardware can succeed depends not only on whether the parameters are close to NVIDIA's, but also on whether the scale of supply, software ecosystem, and economic incentives can drive developers to optimize models for it. Even if the hardware configuration is similar to NVIDIA, if developers cannot obtain sufficient returns from using it, it is also challenging to establish an ecosystem.


The ultimate optimization of the inference platform is not just the model's running speed, but the matching between the model, task, and hardware.


Fireworks Does Not Own Hardware but Seeks to Connect Training and Inference


Faced with AI cloud vendors expanding into the software layer, Fireworks currently has no plans to build its own hardware or make heavy asset purchases for GPUs. Chen Yufei described the company's approach as closer to being a "sublandlord": obtaining computing power from infrastructure providers, with a focus on computing software, inference services, and customized models.


The key difference between Fireworks and typical AI cloud platforms is that most of the traffic on the platform comes not from unmodified base models but from custom models. Chen Yufei believes that the competition for base model inference will become increasingly intense, and profits are more likely to come from delivering exclusive models that can provide better results for customers.


Fireworks integrates training and inference into the same business. The company assists clients in training models and aims for continuous growth in inference traffic after the models are deployed. Unlike service providers that only charge for training or consulting, this revenue structure aligns the platform's long-term usage with the clients' outcomes: as clients generate more value through their models, Fireworks can earn more revenue from ongoing inference.


Some reinforcement learning service providers only offer training or consulting, and their revenue may not be directly tied to the post-deployment performance of the models. Fireworks' business logic is to enhance model performance through training and then reap rewards from subsequent inference traffic. It does not sell one-time training projects but rather the ability for clients' models to continuously generate usage.


The real contenders capable of covering the entire chain are not individual new cloud companies but major cloud services such as AWS, Microsoft Azure, and Google Cloud. They have both computing power and the ability to build reinforcement learning services, inference platforms, and development tools.


Fireworks competes with cloud vendors on one hand while depending on their infrastructure and collaborating with them on the other. Currently, the company has the closest partnership with Azure, allowing customers to use Azure credits to purchase Fireworks services, and it also collaborates with AWS and Google Cloud.


Chen Yufei believes that Fireworks' current inference optimization capability still surpasses similar services from cloud vendors, but a single-point performance is not the most solid barrier. What is more critical is completing the entire machine learning operation cycle from training, deployment, evaluation to inference, providing a seamless experience for developers. Compared to new cloud companies, Fireworks covers a longer chain; compared to cloud giants, it needs to maintain an advantage in specialization and execution speed.


Even if in the future AI can automatically complete kernel, inference, and post-training optimization, understanding customer needs may still be a task that is difficult to fully replace. A system may be able to solve complex mathematical problems, but it may not necessarily be able to spontaneously understand a company's processes, constraints, and judgment criteria. The long-term value of Fireworks ultimately depends on whether customer needs can be translated into models, assessments, and operational products.


Open-Source Model Compression Premium, with Computing Power Still Being the Ultimate Constraint


In the next one to two years, Fireworks's biggest challenge remains computing power. Whether the company can obtain a sufficient amount of computing resources at a reasonable price will directly limit revenue growth. Chen Yufei summarized Fireworks's current state as "constrained by computing power" and assessed that for at least the next year, the overall market may still maintain this state.


This also affects his view on the AI capital expenditures of tech giants. The market usually starts from the return on investment to judge whether companies like Google, Meta, and others are overinvesting; however, for these companies, another risk may be even greater: if AI eventually brings high returns, and they miss out on this round of competition by reducing investment, they may permanently lose their position at the table.


Therefore, even if there is short-term excessive investment, the cost may only be pressure on financial statements and the need to absorb depreciation in the coming years. For tech giants with long-term cash flow and financing capabilities, this is usually not a survival risk; however, missing a key technological cycle may be.


The situation is different for new cloud companies. They need to more rigorously calculate project returns, financing costs, and payback periods. Chen Yufei recommended paying attention to the bond ratings, project delays, and refinancing capabilities of relevant companies. If in the next 3 to 6 months, participants cannot continue borrowing due to credit issues, need debt restructuring, and whether the market still has funds willing to take over will be a key point to test the financing capability of AI infrastructure.


Regarding the market narrative of "open-source models catching up will lower hardware requirements," Chen Yufei is skeptical. In his view, the enhanced ability of open-source models first weakens the premium ability at the model level. If overall AI spending remains the same, the value may instead shift to companies providing computing power and infrastructure.


However, whether AI capital expenditures will ultimately deliver sufficient returns is still an unanswered question. Computing power projects often take years to recoup investments, and the supply-demand landscape a year later is already difficult to predict. The relatively clear signal at this stage is that both Fireworks and many of its customers are still looking for more computing power.


The competition between open-source and closed-source models will not be determined solely by model rankings. As the performance gap continues to narrow, the industry value will depend more on who can define tasks, validate effectiveness, and deploy models into production at sustainable costs.


[Video Link]



Welcome to join the official BlockBeats community:

Telegram Subscription Group: https://t.me/theblockbeats

Telegram Discussion Group: https://t.me/BlockBeats_App

Official Twitter Account: https://twitter.com/BlockBeatsAsia

举报 Correction/Report
Choose Library
Add Library
Cancel
Finish
Add Library
Visible to myself only
Public
Save
Correction/Report
Submit