Classify the failure before choosing the technology
When an assistant misinterprets a product code, department acronym, or internal process name, the immediate reaction is often that the model does not understand the company. In practice, several failures can produce the same bad answer. The system may not have access to the right document, retrieval may return a similar but inapplicable passage, the source documents may conflict, or the model may have the correct evidence but ignore the required response rules.
Start with a reproducible error set. For each example, retain the user question, expected answer, authoritative source, retrieved passages, and generated response. Then label the failure as missing retrieval, wrong retrieval, incorrect interpretation, or noncompliant output. Without this separation, fine-tuning can conceal a retrieval defect, while repeated changes to the vector database cannot repair a behavioral problem.
It also helps to distinguish knowledge from behavior. The current meaning of a part number or department abbreviation is changing knowledge. The requirement to apply a particular decision procedure to a service ticket is learned behavior. Retrieval is usually the better control plane for the first; fine-tuning may help with the second.
Strengthen retrieval when terms change or answers need evidence
If terminology comes from product catalogs, ERP fields, operating procedures, contracts, or a glossary maintained by business teams, retrieval-augmented generation is usually the safer starting point. Updated sources can be indexed without retraining a model, and answers can cite a document, version, and passage. A retrieval layer can also enforce user permissions before any content is sent to the model.
Simply chunking documents and loading embeddings is not a complete RAG design. Enterprise codes are often short, and semantic similarity is unreliable for exact identifiers. The same acronym may mean different things in procurement, finance, and operations. A production system commonly needs keyword search, vector search, metadata filters, and reranking, while preserving enough document structure to interpret the retrieved passage.
- Model terminology as data: Store canonical names, aliases, former names, product codes, owners, and business domains in explicit fields instead of burying them in prose.
- Use hybrid search: Route exact identifiers through keyword or field matching, use vectors for natural-language intent, and rerank the combined candidates.
- Filter by validity and access: Apply effective dates, document status, organization, region, and user role before generation.
- Support abstention: When evidence is missing or authoritative sources conflict, report the gap instead of allowing fluent guessing.
Fine-tune for stable behavior, not a changing encyclopedia
Fine-tuning is more useful when the goal is to reproduce a stable pattern: mapping service descriptions to internal categories, generating maintenance summaries with required fields, applying a consistent professional voice, or selecting among established workflows. It can reduce long prompting and make recurring output structures more reliable. It is much less suitable as the primary store for facts that business teams regularly revise.
Training product names, prices, organizational structures, or the latest procedures into model weights creates an update problem. The model cannot reliably show which version produced an answer, and correcting one obsolete fact is difficult. Training examples also require privacy review, permission checks, deduplication, and careful quality control. If examples contain conflicting policies, fine-tuning may merely reproduce those conflicts more consistently.
- Good candidates: Stable classification schemes, fixed output schemas, domain writing conventions, and decision patterns demonstrated by consistent examples.
- Poor candidates on their own: Frequently revised definitions, live inventory, customer-specific records, and answers requiring traceable citations.
- Check simpler controls first: If clearer prompts, a few examples, or constrained structured output solve the issue, immediate fine-tuning may not justify its data and operational cost.
Most enterprise systems need separation of responsibilities
A robust architecture often assigns retrieval the question of what is currently true and assigns the model the question of how to interpret and present it. A retrieval service might fetch the active definition of an equipment code, the applicable maintenance procedure, and the relevant facility. The model then organizes risks, actions, and open questions according to the company’s required format. Knowledge can change without retraining, while response behavior can be optimized independently.
Four questions usually clarify the choice: How often does the knowledge change? Must the answer cite evidence? Does access depend on the user? Is the failure about factual content or response behavior? If freshness, traceability, or permissions matter, retain retrieval or a live system lookup. Fine-tuning becomes compelling only when correct evidence is already present and the model still fails repeatedly at a stable task.
- Retrieval first: Content changes frequently, answers must be auditable, or users have different access rights.
- Fine-tuning first: Evidence is available, but classification, formatting, or recurring decision behavior remains inconsistent.
- Hybrid design: Terminology changes over time, while the analysis procedure and response contract must remain highly consistent.
Let diagnostic evaluation determine the next investment
Evaluation should measure more than whether an answer appears correct. Build test cases for canonical names, aliases, misspellings, department-specific meanings, obsolete versions, restricted information, missing answers, and conflicting sources. Measure whether the right document was recalled, whether reranking selected the right passage, whether the answer stayed faithful to evidence, whether citations were valid, and whether formatting and abstention rules were followed.
Establish a baseline without fine-tuning, then improve data governance, query rewriting, hybrid retrieval, metadata filters, and prompts in sequence. If evidence retrieval becomes consistently correct but the same behavioral failures remain, fine-tune on a carefully reviewed set of examples and run regression tests against the baseline. This staged approach reveals which layer produced the improvement and reduces the risk of damaging capabilities that already work.
A maintainable enterprise AI system does not try to hide every fact inside a model. It keeps sources, permissions, retrieval, model behavior, and evaluation independently observable and replaceable. Where knowledge is distributed across LINE, ERP, CRM, and cloud platforms, an experienced integration team can help define those boundaries before model choice becomes an expensive distraction.