After years of vendors heralding the arrival of fine-tuned models to save enterprises from the latency, unpredictable costs, and strenuous upgrade lifecycles of cloud-based model APIs – the moment for fine-tuned models has finally arrived. While frontier models are massive, general purpose, and must be run from cloud servers; fine-tuned models (mostly composed of post-trained open-weight models) are smaller, more efficient to run, and can be deployed into devices closer to the edge. Even more crucially they are adapted with companies’ own data to perform better on tasks related to that data. Because these fine-tuned models are based on open-weight models, any enterprise with the wherewithal to build them, can do so, making them more accessible.

The opportunity is not simply that fine-tuned models have become more capable. The more significant change is that enterprises can increasingly use open-weight models as a starting point for specialization, adapting them with proprietary data and domain expertise at a cost that is becoming practical for many more use cases. Historically, the economics and complexity of doing this put customized models out of reach for most enterprises. Better-performing open-weight models and new post-training tooling are now lowering that barrier.

Over the past year we’ve seen a number of new companies (both vendors and enterprises) begin to release their own post-trained custom models (based upon open-weight models) combined with the valuable proprietary data and expertise of these companies. Thomson Reuters introduced Thomson, with training data focused on legal, tax, and regulatory content. Harvey released Tenet, a post-trained model also centering on legal expert data. Travelers announced TravelersLLM to support underwriting analysis and research. Optimizely launched models specialized for marketing applications. Tech Mahindra released Indus, one of several models it is tuning for Hindi and other specific languages.

The rise in companies training and leveraging their own domain-specific fine-tuned foundation models comes in the wake of several others giving up on grander ambitions around more generalized models, realizing that there may be more value in the enterprise with specialization not generalization. Databricks announced it was stopping efforts around its DBRX model after costs continued to run too high, ServiceNow also recently ceased updating their model in favor of supporting other models from 3rd party providers, and even a frontier model provider like Cohere is shifting to provide more directed and specific models.

The development and release pace of post-trained models is still a relative trickle across most enterprises, but there’s significant enablement innovations that could be opening the floodgates. New tooling capabilities, starting with earlier efforts like IBM’s InstructLab to more recent releases like Nvidia’s new Lightning model that seem to significantly improve efficiency of post-training fine-tuning allow these advancements.  And there are numerous other examples such as Axolotl, Unsloth, and Together AI add to this trend.

The opportunity for adoption of fine-tuned foundation models is no longer beyond the horizon for most enterprises. The opportunity will not arrive at the same rate for all enterprises, however. Those with the largest sets of data and intellectual property that either differentiate them in the market or provide a unique cross-cutting view will be best positioned to leverage that into model behavior and business outcomes.

What enterprise technology leaders should do:

  • Assess where the value you can derive from your proprietary data might come from with fine-tuned models. Is it in a specific knowledge domain, or is it tied more to a process domain? What is the relative lifecycle cost of a fine-tuned model versus a RAG or agentic system?
  • Identify specific AI problems that are well suited for a fine-tuned model. Do you need to maintain data or process sovereignty in your AI operations? Do you have employees operating in the field or at remote facilities with limited connectivity to the large cloud-based models?
  • Review your current model routing and failover strategy. Are there any places in your resiliency plan where a self-hosted, task-specific model could help maintain uptime as a failover?

Are you a Forrester client working to understand how fine-tuned models might fit into your AI strategy? Then please reach out to schedule a guidance session.

 

 

 

 

 

 

 

 

 

 

 

Share