Nimit Mehra is the co-founder of Zeo Route Planner, an AI-powered route management platform. He helps companies execute their AI strategy.

​A debate has emerged in the U.S. about how open the next generation of artificial intelligence should be. A recent industry letter supporting open weights model has been backed by more than 270 companies, including Nvidia, Microsoft, OpenAI, Google, Andreessen Horowitz and Y Combinator. Dario Amodei, the CEO of Anthropic, on the other hand, has warned that scaling of open-weight models could pose dangers. As one research report stated, “Once released, they cannot be recalled, and their built-in safeguards can be bypassed through fine-tuning or jailbreaking, posing risks that current governance frameworks are not equipped to address.​”​

What Open-Weight Models Actually Mean​

But what are open-weight models? Large AI models contain billions to trillions of numerical parameters, most of which are weights. During training, these values continuously adjust as the model learns patterns and relationships from data. Think of it crudely as turning interconnected knobs to get the desired signal. Where the knobs stop is the model’s final weights. They are a fundamental part of what the model has learned based on the data.​

In an open-weight model, the developer makes the trained weights available for download under a license. Other developers can then run the model on infrastructure they control or through third-party providers, and, depending on the license and model, fine-tune or otherwise adapt it using their own data. Proprietary models such as Anthropic’s Claude and OpenAI’s closed models do not make their underlying weights publicly available. Instead, developers access them through APIs, applications or managed cloud services, and can customize them only through the capabilities and interfaces the provider makes available. The fundamental difference is who has access to and control over the model weights.​​

Where General-Purpose Models Fall Short​

There is space in the industry for both models to exist and work. The question of which model to choose comes down to a combination of versatility and the complexity of the task, the requirements for control, latency thresholds and the cost for serving the intelligence.​

The current general-purpose models of AI labs are extremely good at writing, creative tasks and coding. Coding in particular is well suited to AI, as the data is easily and abundantly available in digital form, follows explicit syntax and can be executed and tested automatically. Many enterprise workflows, however, work differently. They require proprietary data, internal systems, company-specific workflows and context, often through a sequence of dependent actions. Strong knowledge of a professional domain does not always necessarily translate into reliable execution of end-to-end professional flows.​

Take, for example, accounting. Reasoning models have been passing the CPA exam since 2023, but a recent FinBalance benchmark asked LLMs to create journal entries, reconcile them and then create the balance sheet. Across six LLMs, the best possible accuracy achieved was 46%. The models were able to produce numerically plausible entries but struggled to connect them across the supporting documents and carry them across the complete workflow.​

How Specialization Can Close The Gap​

The answer might lie in specializing the models for professional tasks, rather than expecting a general-purpose model to be the best at all tasks. Recently, researchers fine-tuned the open-weight model Qwen3-VL-2B for reconstructing editable CAD programs from images. Before fine-tuning, the model achieved an accuracy of 6.6%. After fine-tuning, the success rate increased to 82.1%, much higher than the 72.2% of GPT-5.2, which is high for a proprietary frontier model. It shows that on a sufficiently defined task, specialization can compensate for the large gap in general capability.​

The advantages of specialization also extend to cost and can change the economics of serving intelligence at scale. An EnterpriseLab study trained the open-weight model Qwen3-8B on enterprise workflows such as IT, HR, sales and engineering. The specialized 8 billion-parameter model matched the performance of GPT-4o while reducing estimated inference cost by a factor of eight to 10 times.​

So the question to be asked by enterprises should not be which is the best general-purpose model but which is the model that can provide the required accuracy at the lowest possible latency and cost.

Fine-tuning open-weight models makes sense only when the problem is narrowly defined, is repeatable at scale and there are enough proprietary examples of how work is performed with measurable outcomes. The advantages become even more pronounced if there are specific requirements around governance, data location, cost or latency requirements.

For enterprises to be successful at fine-tuning, it requires more than just documents. It requires examples of how experts perform a task, the tools they call, the context they use and the decisions they make in every scenario, especially in difficult cases where judgment is required. This should be treated as an ongoing process rather than episodic. As a company’s workflow, policies and software change, the model would need to be updated as well. Over time, the asset created by the company will not just be the model but also a proprietary repository of training examples, expert corrections, evaluations and workflow knowledge.​

The Full Cost Of Owning An AI Model​

However, companies should not only look at cost of inference to evaluate open-weight models. The economics should be evaluated on the full cost of ownership, including compute infrastructure, data/cloud storage cost, maintenance and security. Additionally, with open-weight models, the responsibility of security, governance and access now lies with the company rather than model provider. In some cases, taking into consideration these responsibilities and costs, a general model might still be beneficial, while in other cases, open-weight fine-tuned models might work best at scale.​

Open-weight models will not replace general-purpose models, nor do they make sense for every enterprise workflow. But as AI gets embedded more and more in our lives, the decision that companies need to make is which parts of AI they comfortable renting indefinitely and which parts should become proprietary organizational assets.​

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Share.
Leave A Reply

Exit mobile version