Treat consistency as a data contract
Enterprise assistants often read Chinese policies, English technical manuals, ERP fields, and service knowledge bases in the same session. If the only instruction is to answer in the user's language, one concept may receive several plausible translations. A purchase order might appear as a full English term in one response, an abbreviation in another, and a loosely translated Chinese label in a third. Each sentence can sound natural while search, reporting, and downstream automation quietly lose consistency.
The engineering response is a machine-readable terminology contract. Give each concept a stable identifier, approved Traditional Chinese and English labels, permitted abbreviations, rejected variants, scope, owner, and authoritative source. APIs, retrieval indexes, and evaluations should use the stable identifier; the presentation layer selects the approved label for the current language. This moves terminology from a model preference to governed application data.
Apply terminology controls throughout the pipeline
A glossary appended to a long system prompt is easy to overlook, especially after several conversation turns or large retrieval results. Use terminology data at four points instead. Query normalization maps aliases to concept identifiers. Retrieval expands the query with approved bilingual terms and abbreviations. Generation receives only the entries relevant to the current answer. Post-generation validation detects prohibited translations, unexplained abbreviations, and inconsistent names.
Source language should also be preserved. Legal titles, product features, API names, and ERP fields are often safer in their original form. Other domain terms may use an approved Chinese label followed by the English term on first mention. Define the choice by content type. Showing both languages improves verification but adds visual noise; showing only the response language reads better but makes the original document harder to locate.
- Fixed names: Preserve company, product, API, table, and field names; add an explanation only when needed.
- Controlled translations: Use approved bilingual mappings for domain terminology and define when an abbreviation may appear.
- Contextual prose: Allow natural phrasing, but never alter numbers, conditions, responsible parties, or negation.
- Unknown terms: Retain the source wording and flag it for review instead of inventing an official translation.
Make citations point to evidence
Citation consistency has three parts: a claim must be supported by its cited passage, both language versions should rely on evidence of comparable authority, and the user must be able to reopen the referenced content. Asking a model to add document titles at the end of a response is insufficient. It may choose a related passage that does not support the claim or produce a title that resembles a real document. Citation identifiers should come from retrieval, and the model should only select from identifiers supplied by the application.
During ingestion, retain the document identifier, version, section, page or paragraph locator, language, and access policy. If Chinese and English files are official versions of the same document, connect them through a shared document family and version relationship. If one is merely a summary, label it as such. Place citations near factual claims, and keep an immutable source key behind the display title so renaming a file does not break traceability.
- Completeness: Key conclusions, limitations, dates, and required actions have supporting evidence.
- Entailment: The cited passage directly supports the adjacent claim, not merely the general topic.
- Cross-language parity: Switching languages does not replace an authoritative policy with a weaker secondary summary.
- Access parity: The response and citation preview enforce the same document permissions.
Control drift with reproducible evaluation
Bilingual quality cannot be established by reading a few polished examples. The evaluation set should include Chinese questions against Chinese sources, English against English, Chinese against English sources, English against Chinese sources, and language changes within one conversation. Besides factual correctness, test approved terminology, prohibited variants, numerical fidelity, source identifiers, citation coverage, and permission behavior. Deterministic rules work well for exact constraints; human or model-based review is still useful for meaning and readability.
Version the glossary, prompts, embedding model, chunking policy, and source documents. Run the same representative questions after each change so the team can identify which layer caused an improvement or regression. Production traces should record the model version, glossary version, retrieved sources, and citation mapping while avoiding unnecessary sensitive content. When a user reports a problem, engineers can then reproduce the answer instead of guessing why the wording changed.
- Assign authority: Product owners govern voice, domain experts approve terms, and document owners confirm source validity.
- Review exceptions: Route new terminology to a pending queue instead of turning it immediately into a global rule.
- Separate failure modes: Track mistranslation, term drift, citation mismatch, stale sources, and access leakage independently.
- Provide an escalation path: When evidence is missing or versions conflict, the assistant should state the limitation and request confirmation.
A mature bilingual assistant does not need identical sentences in both languages. It needs equivalent concepts, evidence strength, and operational consequences. When the assistant also connects to LINE, CRM, ERP, or cloud knowledge stores, terminology and citation rules should enter the integration data model early instead of being patched separately into every channel.