Choosing Between Domain-Specific and General Purpose AI Models
When specialized models outperform general-purpose ones

Source: Father_Studio/stock.adobe.com
As AI tools spread across industries, a clear distinction has emerged between general-purpose models built for broad linguistic and reasoning tasks and domain-specific models optimized for narrow, high-stakes applications. This blog defines what makes a model "domain-specific" and examines how targeted training data, fine-tuning approaches, and evaluation criteria differ from their general counterparts. Use cases in law, medicine, and finance illustrate where specialization delivers measurable advantages and where it introduces new constraints related to flexibility and maintenance costs.
What Makes a Model Domain-Specific
General-purpose models are trained on a broad mix of internet text, code, and books, then tuned to follow instructions across a wide range of tasks. A model becomes domain-specific when it is fine-tuned or trained from scratch on domain-relevant datasets typically compiled by subject-matter experts. Domain-specific models are also evaluated using benchmarks designed for the specific field rather than the broader tasks used to assess general-purpose models.
Fine-tuning adapts a model with general language and reasoning capabilities by exposing it to more specialized data and adjusting its parameters, for example, via Low-Rank Adaptation (LoRA) to shift its behavior toward a specific domain. This approach works well when the base model already contains most of the relevant knowledge, and what is missing is domain-specific vocabulary or task-specific judgment. By contrast, training a model from scratch requires a custom mix of domain-specific and general text so the model learns language structure, general knowledge, and domain knowledge simultaneously from the training data. The dataset must therefore be large and diverse enough to teach the model to write coherently while also developing the required specialization. This approach is appropriate when a domain's vocabulary, structure, or proprietary data is sufficiently distinct that no existing model has adequate prior knowledge of it.
There are also approaches that fall between these two, as illustrated by the real-world examples that follow. Regardless of the approach, the decision to use a domain-specific model may be driven by the need to host AI systems internally due to data privacy or regulatory constraints. The most capable general-purpose models may outperform domain-specific alternatives, but they are not always easy to use. Many are too large to host internally, making self-hosting impractical. In some cases, self-hosting those models is not possible because their parameters are proprietary and not publicly released. In other words, the model cannot be run outside the provider's infrastructure. In these situations, smaller general-purpose models fine-tuned for a specific domain can serve as a practical alternative. While self-hosting comes with added costs for infrastructure and for building the right datasets for training and evaluation, it can give companies complete control over the deployed models and the data that runs through them.
Real World Examples
BloombergGPT, a finance-focused large language model (LLM), demonstrates the from-scratch approach. Trained on a 700-billion-token mix of general and domain-specific data, it beat similarly sized open models on financial tasks by wide margins. For example, it scored 43.41 percent versus 30.06 percent on ConvFinQA, a benchmark that tests multi-step numerical reasoning in conversational finance question answering, and 75.07 percent versus 50.59 percent on FiQA SA, a financial sentiment analysis benchmark, while remaining competitive on general-purpose benchmarks.[1] FinGPT took a lower-cost approach, fine-tuning an existing open model with LoRA, and reportedly matched or outperformed BloombergGPT on several tasks at considerably lower training cost.[2] A separate evaluation later found that BloombergGPT underperformed newer general-purpose models released after it, suggesting that a domain-specific model's advantage can narrow as general-purpose models continue to improve.[3]
Medicine illustrates a related dynamic. Med-PaLM 2, fine-tuned specifically on medical question-answering data, achieved 86.5 percent accuracy on a dataset of questions similar to those on the United States Medical Licensing Examination (USMLE), making it the first model to reach expert-level performance on that benchmark.[4] GPT-4, a general-purpose model with no medical fine-tuning, then reached 90.2 percent on the same benchmark using prompting techniques alone. Microsoft researchers found that a sufficiently capable general-purpose model, when prompted effectively, could match or exceed the performance of specialized medical fine-tuning across several benchmarks.[5] Their findings show that a sufficiently large general-purpose model can eliminate most of the performance advantage that domain-specific fine-tuning was designed to provide, at least for knowledge already well represented in publicly available medical literature.
Law presents a third approach. Harvey, used across major law firms, does not train a single model from scratch. Instead, it routes each request through a combination of fine-tuned and general-purpose models, supplemented by retrieval systems that draw on case law, filings, and firm-specific documents.[6] In an independent benchmark, Harvey's tools outperformed a human lawyer baseline on most tasks tested, including document question-answering and document extraction.[7] Harvey's design shows that domain specialization can be achieved through an orchestration layer of routing, retrieval, and fine-tuning built around general-purpose models, rather than through a single specialized model.
The Limits of Domain-Specific Training
Two of the three examples above follow the same pattern. BloombergGPT's advantage over open models narrowed once newer general-purpose models became available, and Med-PaLM 2's advantage over an untuned GPT-4 disappeared once prompting techniques improved. Domain-specific training can produce an accuracy advantage at a given point in time, though not always a lasting one, since general-purpose models continue to incorporate more of the public knowledge that previously required specialization to access. In addition, a domain-specific model will still require retraining or fine-tuning whenever its field changes, such as when new case law is introduced or treatment guidelines are updated. Harvey's approach, which routes requests through a combination of general-purpose and domain-specific fine-tuned models selected according to the task, with retrieval systems providing access to relevant documents, indicates the direction several fields are likely to move toward.
Conclusion
Domain-specific and general-purpose models differ in how they are built and evaluated. Domain-specific models are created by fine-tuning an existing model or training one from scratch to better align with a field's vocabulary and task-specific requirements. They also require domain-specific evaluation methods, tailored to the domain, often requiring practitioners in the field to assess outputs and construct evaluation sets, rather than relying on generic benchmarks.
In practice, domain-specific models can sometimes achieve higher accuracy than general-purpose alternatives. However, as general-purpose models improve and prompting techniques advance, that performance gap tends to narrow. Even so, there are other reasons for choosing a domain-specific model. Data privacy requirements, regulatory obligations, or other operational constraints may prevent organizations from using externally hosted AI services. In these situations, developing an internal domain-specific model may be justified, even if a larger general-purpose model would otherwise perform better.
Choosing that path, however, also means accepting responsibility for keeping the domain-specific system current as the field evolves. The decision between general-purpose and domain-specific, therefore, comes down to identifying which constraints apply to a given situation before committing to either approach.
[1]https://arxiv.org/abs/2303.17564
[2]https://arxiv.org/pdf/2505.19819
[3]https://arxiv.org/pdf/2401.14777
[4]https://www.nature.com/articles/s41591-024-03423-7
[5]https://www.microsoft.com/en-us/research/blog/advances-in-run-time-strategies-for-next-generation-foundation-models/
[6]https://help.harvey.ai/articles/what-ai-models-does-harvey-use
[7]https://www.vals.ai/industry-reports/vlair-2-27-25