Glossary Operations · Published 2026-06-26

Caption glossary maintenance workflow: term submissions, product rebrandings, acquisition onboarding, version control, and LMS catalog synchronisation

A caption glossary built at programme launch degrades the moment your first product ships a new feature name. The glossary architecture post covers how to structure a glossary — term types, phonetic pairs, confidence scoring, domain tagging. This post covers what happens after structure: the operational questions that determine whether the glossary stays accurate at month 6, month 18, and month 36. Who can submit a new term? Who approves it? What happens when the product team announces a rebranding at 09:00 on a Monday and the first training module that uses the new name is being recorded at 13:00 the same day? What do you do with the acquired company’s 340-term product vocabulary the week the acquisition closes? What happens when you deprecate a product name but 47 training modules from 2023 still use it? These are maintenance questions, not architecture questions, and the gap between the two is where most production caption glossaries quietly fail. This post gives you the operational framework to close that gap.

TL;DR

Five things to know before you build a glossary maintenance workflow:

  1. Maintenance is a distinct programme function from setup. Building the glossary is a project. Maintaining it is an ongoing operational function that requires assigned ownership, defined workflows, and a quarterly review cadence. Teams that treat maintenance as “updating the glossary when someone notices an error” accumulate drift at a rate that erases the accuracy gains from the initial glossary build within six months.
  2. Product rebrandings require a glossary update on announce day, not the following week. The first training module recorded after a product name change embeds the new name in audio. If the glossary still contains only the old name, the ASR model transcribes the new name incorrectly. The fix — a pre-announcement glossary update under NDA — is simple, but it requires a standing workflow that connects the product marketing team to the glossary custodian before the announcement goes out.
  3. Term deprecation is not the same as deletion. Removing a deprecated product name from the glossary breaks caption accuracy on archive content that was produced when the name was current. Deprecated entries must stay in the glossary as inactive entries, flagged so the caption model does not preferentially substitute them in new content but still recognises them in content where they were the correct term at time of production.
  4. Every caption job should log which glossary version it used. This is the DCMP audit trail requirement: if an accessibility complaint or OCR investigation asks you to demonstrate that your captions were produced with the correct vocabulary for the content type and production date, the version log is your evidence. Semantic versioning (v2.3.1) tied to release dates makes the log queryable.
  5. LMS course archive events should trigger a glossary review, not a glossary deletion. When a course is retired from the LMS catalog, its glossary terms should be reviewed — not automatically deleted. Some terms serve multiple courses; deletion of a retired course’s unique terms may leave live captions in companion courses without vocabulary coverage.

Why glossary maintenance is a distinct programme function

The glossary architecture post covers the design decisions that determine whether a caption glossary produces reliable results at launch: which term types to include (product names, SDK symbols, acronyms, regulatory vocabulary, medical terminology), how to structure phonetic variants, how confidence scoring affects substitution behaviour, and how domain tagging routes terms to the right content types. Those are one-time structural decisions.

Maintenance is not a continuation of architecture. It is a different activity with different inputs, different ownership requirements, and different failure modes. Architecture fails when the structure is wrong. Maintenance fails when the structure is right but the content goes stale. And because most organisations build a glossary during onboarding, celebrate the accuracy improvement, and then assign no one to keep it current, maintenance failure is far more common than architecture failure in production caption programmes.

Three dynamics drive glossary drift in a production environment:

Vocabulary evolves faster than a self-updating glossary can track. A 50-employee SaaS company releases 4–6 named features per quarter. A 200-employee company may release 15–20. Each release introduces new product names, updated technical terminology, revised acronyms, and changed feature-set vocabulary. Without a standing workflow to capture these releases and translate them into glossary updates, the glossary lags product reality by one to three quarters within the first year of operation. The glossary-biased decoding post shows what this lag looks like in practice: engineering onboarding content produced three quarters after a major SDK rename produces substitution errors on the new SDK name at the same rate as it would without a glossary, because the glossary entry still uses the deprecated name and the model has no recognition path for the current one.

The people who know what terms belong in the glossary are not the people who maintain it. The term custodian who owns the glossary day-to-day is typically an L&D producer or an accessibility programme manager. The people who know that a new feature shipped last Tuesday, or that the legal team has changed the preferred acronym for a regulatory programme, are in product marketing, engineering, or legal. Without a formal submission channel that connects those domain experts to the glossary, updates flow through informal channels — a Slack message, a comment in a shared doc, a note from a vendor reviewer — or they don’t flow at all. The captioning governance policy post covers how to embed the glossary submission channel into the organisation’s content production policy; this post covers the operational mechanics of that channel.

Archive content and current content use the same glossary, but they need different things from it. A training module produced in 2022 that refers to a product that has since been renamed needs the glossary to recognise the 2022 name as correct for that content. A training module being produced today needs the glossary to recognise the 2026 name as correct. A glossary that serves both simultaneously must have a concept of term lifecycle — active, deprecated, retired, and with version tracking that logs which state each term was in at the time of each caption job. Organisations without this lifecycle concept delete old names from the glossary when they add new ones, which corrects new content while breaking archive content.

Maintenance as a programme function means: assigned ownership (someone is accountable for the glossary state), defined workflows (there is a specific path for submitting, approving, and deploying a term update), a quarterly review cadence (the glossary is audited on a schedule rather than reactively), and a version history that serves both operational and compliance purposes. Each of these is covered in detail below.

Ownership model: the three-role RACI

A production caption glossary requires three distinct roles. These can be filled by three individuals, two individuals with one playing a dual role, or in smaller organisations a single person who performs all three functions at different points in the workflow. What cannot work is assigning all three functions to an implicit group (“the L&D team”) with no individual accountable for any of them. The building a caption compliance programme post covers the broader programme role design; the glossary RACI is a subset of that.

Role 1: Term Custodian

The Term Custodian is the operational owner of the glossary. This role is responsible for day-to-day submissions, the submission queue, the approval workflow, version releases, and the quarterly audit. In most L&D teams, this is the accessibility programme manager, a senior L&D producer, or the captioning programme lead. The Term Custodian does not need to be a subject-matter expert in the content domains; they need to be an expert in the glossary system and the maintenance workflow.

The Term Custodian’s specific responsibilities:

Role 2: Domain Authority

The Domain Authority is the subject-matter expert who confirms that a submitted term is correct, is used as described, and should be applied to the content types specified. This role is typically filled by someone in product marketing (for product terms), engineering (for SDK and technical terms), legal or compliance (for regulatory vocabulary), or clinical operations (for medical terminology in healthcare organisations). The Domain Authority is not a glossary operations role; it is a knowledge confirmation role invoked when the Term Custodian cannot independently verify whether a submitted term is accurate.

Not every term submission requires Domain Authority review. Routine additions — a new acronym that appears in a support document, a minor feature name update that is publicly documented — can be approved by the Term Custodian alone. Domain Authority review is required for: new product names before public announcement (to confirm the name is final and the spelling is correct), regulatory vocabulary changes (to confirm the new term reflects a genuine regulatory update, not an informal alternative), medical terminology (any term with clinical accuracy implications), and terms where the submission conflicts with an existing entry.

Role 3: Vendor Bridge

The Vendor Bridge is the interface between the approved term list and the caption model. This role translates approved terms into the format the captioning service requires: phonetic pronunciation variants for ASR, confidence scores, context restrictions by content type, and the actual upload or API call that deploys the term to the live model. In many organisations, the Vendor Bridge is not an internal role at all — it is a function performed by the captioning vendor’s customer success or technical team on receipt of approved terms from the Term Custodian. However, the internal Term Custodian remains accountable for verifying that the deployment happened correctly and that the deployed term behaves as expected in a test transcription before the version is released to production.

The RACI below shows how these three roles interact with each stage of the maintenance workflow:

Workflow stage Term Custodian Domain Authority Vendor Bridge
Term submission received Accountable Informed (if relevant domain) Informed
Triage: priority assignment Accountable Consulted (urgent terms)
Domain Authority review (when required) Responsible (routing) Accountable
Phonetic variant generation Responsible (coordination) Consulted (pronunciation confirmation) Accountable
Model deployment Responsible (verification) Accountable
Version release Accountable Informed Informed
Production verification (test transcription) Accountable Consulted (if term is complex) Responsible
Deprecation decision Accountable Consulted Informed
Quarterly audit Accountable Consulted (per domain) Responsible (accuracy data)

Term submission workflow: who submits, how, and when

The question of who can submit a term to the glossary has a surprising amount of operational consequence. Two patterns exist in production L&D programmes, and they produce different maintenance dynamics.

Open submission: any team member can submit

In an open-submission model, any L&D producer, instructional designer, subject-matter expert, or content reviewer can submit a term. The Term Custodian receives all submissions, triages them, routes for Domain Authority review if required, and approves or rejects.

The advantage of open submission is coverage: the people most likely to notice that a term is missing or wrong are the people actively producing content. A producer who has just watched a transcript where “Azure AD B2C” was transcribed as “Azure ADB to see” knows exactly what term is missing and can submit it in the moment.

The disadvantage is volume management. An open-submission channel without triage discipline produces a queue with a mix of genuinely high-priority updates (the product name that shipped yesterday), low-priority additions (a synonym that has one occurrence in the entire library), and duplicate submissions (five people submitted the same term from five different modules). The Term Custodian must deduplicate and triage before processing, which adds overhead proportional to programme size.

Nominated custodian submission: designated term contributors only

In a nominated-custodian model, term submissions come from a named set of contributors: one or two people per content domain who are authorised to add terms to the queue. All other team members route their term requests through those nominated contributors.

The advantage is queue quality: nominated contributors know the submission process and submit well-formed requests that require less triage. The disadvantage is coverage gaps: if a nominated contributor for the engineering domain is on leave when a major SDK update ships, the glossary does not get updated until they return.

For most L&D teams up to 50 people, open submission with a structured form is preferable. Above 50 people, the nominated-custodian model with one contributor per content vertical (compliance, technical, medical, general) reduces queue overhead significantly. In either case, the submission form fields are the same:

Submission form fields

Field Required Purpose
Term (exact form) Yes The term as it should appear in captions. Include capitalisation, spacing, and punctuation exactly.
Phonetic pronunciation Recommended How the term sounds when spoken, especially for acronyms (HIPAA = “HIP-uh”), initialisms (AWS = “A-W-S”), or unusual proper nouns. Without this, the Vendor Bridge must research or guess.
Term type Yes Product name, feature name, SDK/API name, acronym, regulatory term, medical term, person name, organisation name, or other.
Content domain(s) Yes Where this term appears: compliance, technical, medical, soft-skills, executive, or all. Determines which content-type routing rules apply in the model.
Example sentence Recommended A sentence from the actual content where this term appears. The most useful single input for the Vendor Bridge phonetic and confidence-scoring step.
Priority Yes Urgent (needed within 48 hours — active production is blocked), standard (next weekly batch), or low (accumulate for quarterly review).
Submitter and date Yes For audit trail and follow-up.
Linked course or module Recommended Which course triggered the submission. Used for LMS catalog synchronisation and deprecation decisions later.

Submission cadence: batch vs continuous

Standard submissions are processed in weekly batches. The Term Custodian opens the queue once per week, deduplicates, triages, routes for SME review if required, collects approvals, and packages the batch for the Vendor Bridge step. This batching keeps the workflow manageable and aligns version releases to predictable intervals (one minor version per week, for example).

Urgent submissions bypass the batch cadence. Any term flagged as urgent triggers an out-of-cycle workflow: the Term Custodian reviews within 4 hours of receipt, routes for any required Domain Authority confirmation within the same business day, and coordinates an expedited Vendor Bridge deployment targeting a 48-hour turnaround from submission to live model update. The 48-hour SLA for urgent terms is structural: a training module recorded after a product announcement with a glossary that doesn’t know the new name will produce incorrect captions that require re-delivery, which costs more than the 48-hour expedited update. The caption feedback loop post quantifies the accuracy compounding effect of timely glossary updates; the same logic applies in reverse to timely updates that prevent accuracy degradation from new vocabulary.

Approval chain: single-step vs two-step, and when SME review is required

The approval chain for a glossary term governs whether a submitted term is added to the production model or held for further review. An approval chain that is too light (Term Custodian approves everything unilaterally) produces a glossary with incorrect terms and phonetics. An approval chain that is too heavy (every term requires three rounds of SME sign-off) produces a backlog that means the glossary is always three months behind product reality. The right calibration depends on the term type.

Single-step approval (Term Custodian only)

These term types can be approved by the Term Custodian without Domain Authority review:

For these terms, the Term Custodian verifies the correct form from a public source (the product documentation, the regulatory citation, the company’s press release), records the source in the submission log, and routes the term directly to the Vendor Bridge step.

Two-step approval (Term Custodian routes to Domain Authority)

These term types require Domain Authority confirmation before the Term Custodian approves:

Vendor Bridge step: phonetic variants and confidence scoring

Once a term clears the approval chain, the Vendor Bridge translates it into the format the caption model requires. This step has three outputs:

Phonetic variants: The spoken forms of the term that the ASR model should recognise and map to the written form. A product name like “Accenture” typically needs one variant (“AK-sen-cher”). An acronym like “CISA” may need two — the spelled-out form (“C-I-S-A”) and the word-form pronunciation (“SEE-suh”) depending on how speakers use it in the organisation’s content. The proper noun failure modes post covers the phonetic categories that produce the highest error rates; the same taxonomy applies when generating phonetic variants for the Vendor Bridge.

Confidence score: The weight the model assigns to this term relative to competing transcription candidates. A high-confidence term (score 0.9+) will be substituted for any phonetically similar output that doesn’t match a higher-priority term. A low-confidence term (score 0.5–0.7) allows common words that sound similar to win in general content. Confidence scoring should be calibrated to the term’s frequency in the relevant content type and its phonetic uniqueness. Setting everything to maximum confidence creates false positives on common words that happen to sound similar to the glossary term. The glossary vs prompting vs fine-tuning post covers how confidence score calibration interacts with the underlying ASR architecture.

Context restriction: Whether the term should be active for all content types or restricted to specific domains. A medical drug name with a high confidence score should be restricted to content tagged as medical or compliance; activating it across all content types may produce false substitutions in general-skills content where the phonetic sequence occurs as common words.

After the Vendor Bridge deploys the term, the Term Custodian runs a production verification: a short test transcription using audio that contains the target term in context, to confirm the deployed entry produces the expected output. If the verification reveals unexpected behaviour (over-suppression of a similar term, false positive in adjacent vocabulary), the term is held out of the version release and returned to the Vendor Bridge for adjustment.

Product rebranding: the announce-day problem

Product rebrandings create the single most time-sensitive maintenance event in a caption glossary lifecycle. The problem is structural: the first training module recorded after an announcement uses the new product name in audio. If the glossary was not updated before that recording session, the ASR model transcribes the new name incorrectly. The incorrect caption is then embedded in the module, uploaded to the LMS, and may be used in training sessions before the error is caught — if it is caught at all, because reviewers who are also adjusting to the new name may not notice that the captions use a wrong transcription rather than the old name.

Why announce day is the critical moment

The glossary update lag is not primarily a technical problem — deploying a new term to the caption model takes hours, not days. It is an information-flow problem: the L&D team learns about the rebranding at the same time as the public, after the announcement. By that point, content production has begun. The gap is typically 0–72 hours between announcement and first affected training content being recorded, and the captioning correction cycle (return to vendor, re-deliver, re-upload to LMS) takes 3–5 business days. An organisation that does not have a pre-announcement update workflow pays for the lag in re-delivery costs and in captions that are live in the LMS carrying the incorrect transcription during the correction window.

The pre-announcement workflow

The pre-announcement workflow requires a standing relationship between the product marketing team (or whoever manages brand vocabulary) and the Term Custodian. The mechanics are simple:

  1. The product marketing team notifies the Term Custodian of the rebranding under NDA, typically 1–2 weeks before announcement. The notification includes: the new product name (final, approved form), the old product name being replaced, the announcement date, and any alternative name forms that will appear in public documentation (short forms, abbreviated forms, “formerly known as” forms).
  2. The Term Custodian submits the new name through the standard urgent approval workflow, with a scheduled deployment date of announcement day −1 or announcement morning before business hours, depending on the organisation’s content production schedule.
  3. The Domain Authority confirms the final spelling (product marketing owns this). The Vendor Bridge pre-stages the deployment for activation at the scheduled time.
  4. The old name is flagged for deprecation on announcement day, with the deprecation entry created simultaneously with the new name entry (see the deprecation section below).

This workflow depends on the Term Custodian being included in the pre-announcement communication. Including captioning glossary maintenance in the organisation’s product launch checklist is the structural fix that makes this routine rather than reactive. The captioning governance policy post covers how to embed this requirement in the policy document that product marketing, engineering, and L&D all sign off on.

In-flight content at the time of rebranding

Content that was in production at the time of the announcement presents a branching decision: modules that have been recorded but not captioned will pick up the new glossary entry automatically if they are submitted for captioning after the deployment. Modules that have already been captioned using the old glossary need to be assessed for re-delivery.

The re-delivery decision should be based on whether the old name appears in the caption track and whether the old name is a recognisable substitution error (the transcript reads the old name correctly) or a mistranscription (the transcript has a phonetic approximation of the new name using the old model). In the first case, the captions are technically correct for the content as it was produced — the speaker said the old name, the caption reads the old name, and the content may be updated or retired as part of the rebranding rollout. In the second case, the captions are incorrect for the content as produced, and re-delivery is required regardless of rebranding timing.

Announcing mid-series

A rebranding that occurs in the middle of a multi-module training series creates a consistency problem: earlier modules in the series use the old name correctly; later modules should use the new name. The standard approach is not to re-caption the earlier modules (the old name is historically correct for content produced before the announcement) but to add a series-level note in the LMS course description that modules produced before [announcement date] use the previous product name. The glossary entry for the old name should carry a deprecation date that corresponds to the announcement, so the entry’s lifecycle is documented in the audit trail.

Acquisition vocabulary onboarding

An acquisition creates a sudden, large-scale glossary update requirement: the acquired company’s product names, technical terms, internal acronyms, and domain vocabulary enter the primary organisation’s training content catalogue from the moment integration activities begin. Without a structured onboarding process, the glossary is missing all of this vocabulary when the first integration training modules are produced, and the accuracy degradation is indistinguishable from the kind documented in the QA methodology post — proper nouns substituted with phonetic approximations, acronyms spelled out incorrectly, product names mangled.

Priority tiering for acquired vocabulary

Acquisitions involve too many terms to onboard at once at urgent priority. A structured tiering approach manages the volume while ensuring the most critical vocabulary is available before first use:

Priority tier Term types Target deployment Source
Tier 1 (Urgent) The acquired company’s primary product name(s), flagship brand names, and any names that appear in the acquisition announcement Announcement day Acquisition announcement document
Tier 2 (Within 1 week) Product suite names, platform names, major feature names, C-suite and key leadership names if they will appear in integration training content Week 1 of integration activities Acquired company’s product documentation, press kit
Tier 3 (Within 1 month) Technical vocabulary, SDK names, API names, internal acronyms, regulatory vocabulary specific to the acquired company’s vertical Before first technical training modules are recorded Acquired company’s internal glossary (if available), engineering documentation
Tier 4 (Quarterly batch) Historical product names (retired but still referenced in archive content), informal vocabulary, regional variations Quarterly audit following acquisition close Content audit of acquired company’s training library

The most valuable source for Tier 2 and 3 vocabulary is the acquired company’s own internal glossary, if one exists. Many technology companies maintain a vocabulary guide for marketing and documentation consistency; this guide, adapted for the captioning use case, can reduce the Term Custodian’s research workload by 60–80% for a typical SaaS acquisition. Even a brand style guide with a product name section provides enough to bootstrap the Tier 2 vocabulary list.

Conflict resolution between acquiring and acquired vocabulary

Vocabulary conflicts are common in acquisitions. The acquiring company and the acquired company may have different internal names for equivalent concepts (“workspace” vs “project” for the same construct), or they may have overlapping feature names for different features (“Flows” as both an acquired company workflow product and an existing acquiring company automation feature).

The conflict resolution process for glossary purposes is distinct from the product integration decision. The glossary must document both terms accurately for the content in which each term is used, even if the product roadmap is still determining which name survives the integration. The approach is to maintain both conflicting terms as active entries with content-domain context restrictions, and to update the restriction logic when the product team makes the final naming decision. Premature deletion of one conflicting term causes caption errors on content that was produced using that term correctly.

When the acquired company had no glossary

Many acquired companies, especially companies with fewer than 50 employees or companies without an established L&D function, will not have an internal vocabulary guide. In this case, the Term Custodian must bootstrap the acquisition vocabulary from first principles: reviewing the acquired company’s public product documentation, press releases, and feature documentation to extract proper nouns, and conducting a brief review session with an acquisition team member who can confirm the internal vocabulary conventions. A 30-minute terminology session with someone from the acquired company’s product or marketing team typically yields a reliable Tier 1 and Tier 2 vocabulary list that prevents the most visible captioning failures during the first wave of integration training content.

Term deprecation without deletion

Deprecation is the most operationally misunderstood part of glossary maintenance. The intuitive action when a product is renamed is to update the glossary entry: change the old name to the new name, and the next time the caption model sees audio that previously produced the old name, it will produce the new name. The problem is that this intuitive action is incorrect for archive content.

The archive content problem

If a training module produced in 2023 uses the term “Horizon” correctly (because that was the product name in 2023) and the product was renamed to “Apex” in 2025, the 2023 module should continue to have “Horizon” in its captions. A Deaf learner accessing the 2023 module in 2026 should see “Horizon” in the captions, which is what was said in the audio. If the glossary was updated to replace “Horizon” with “Apex” and the 2023 module was re-captioned using the updated glossary, the captions would now say “Apex” for audio that says “Horizon”, creating a new caption error rather than fixing the rebranding.

The correct model is deprecation-without-deletion: the old term remains in the glossary as an inactive entry that is still recognised for existing content where it was the correct term, but is not applied as a first-choice substitution candidate in new content where the new term should appear instead.

Deprecated entry format

A deprecated glossary entry should contain the following fields, in addition to the standard term metadata:

Field Value Purpose
Status deprecated Flags the entry as inactive for new content production decisions
Deprecation date The date the term was deprecated (announcement day for rebrandings, acquisition close date for acquired company terms) Allows the version log to determine which terms were active at any given production date
Superseded by The new term that replaces this one Creates a navigable replacement chain; used by the Vendor Bridge when determining how to handle phonetic overlaps
Archive content scope Date range of content that correctly uses this term (e.g., “2021-03-01 to 2025-08-14”) Defines which content jobs should preserve the deprecated term in captions vs receive the new term
Sunset date Optional date when the entry is eligible for full retirement (see below) Triggers quarterly review reminder when the sunset date is reached

How deprecated entries behave in the caption model

The technical implementation of deprecation varies by captioning service, but the logical behaviour should be:

The sunset date pattern

Deprecated entries accumulate over time. An organisation with a 5-year-old caption programme and three product rebrandings may have 40–60 deprecated entries that are still technically needed for archive content. At some point, that archive content will be retired from the LMS, and the deprecated entries can be fully removed without affecting any live content. The sunset date pattern handles this:

  1. When a term is deprecated, the Term Custodian assigns a provisional sunset date based on the organisation’s content retention policy. For content with a 5-year retention period, the sunset date is 5 years after the deprecation date.
  2. At the quarterly audit, any entries within 3 months of their sunset date are reviewed: is the archive content that uses this term still in the LMS catalog? If yes, the sunset date is extended. If no, the entry is retired.
  3. Retired entries are moved to a separate archive log (not deleted from the records), with the retirement date and the name of the archive content that was their last use. This preserves the audit trail while removing the entry from active glossary maintenance.

A glossary entry is only permanently deleted when: the archive content that used it is confirmed retired from all LMS catalogs, the entry has been in the archived state for a full audit cycle with no re-activation requests, and the Term Custodian and Domain Authority have both signed off on the deletion. This two-person sign-off requirement prevents single-actor accidents that remove entries with live dependencies.

Version control: tagging, logging, and rollback

A production caption glossary without version control is an accuracy black box. When an accessibility complaint arrives 18 months after a module was captioned, or when the annual programme review asks why accuracy on medical content declined in Q3, the organisation needs to answer the question: what glossary state was in use when that content was captioned? Without version control, that question cannot be answered.

Semantic versioning for glossary releases

The semantic versioning convention used in software development applies directly to glossary management:

The version tag is assigned by the Term Custodian at the time of release and recorded in the version log alongside: the release date, the number of terms added, the number of terms deprecated, the number of terms modified (phonetics or confidence scores), and the specific terms changed (for minor and patch releases). For major releases, the log includes a summary of the structural changes and the re-testing outcomes.

Per-job version logging

Every caption job submitted to the production model should have its glossary version recorded in the job metadata. The minimum required fields are:

This log is the DCMP audit trail evidence. If an OCR accessibility complaint names a specific module, the Term Custodian can pull the job log, identify the glossary version active at captioning time, and produce the glossary state for that version from the version history. This demonstrates that the vocabulary appropriate to that content type and production date was in use, which is the correct response to a methodology challenge in a compliance investigation.

The QA methodology post covers how version log entries connect to the accuracy scoring records kept by the QA function. The two logs together — glossary version at production time, accuracy score at QA time — create a complete chain of evidence from vocabulary state to accuracy outcome. This chain is what makes a caption programme audit-ready, as distinct from merely having captions.

Rollback procedure

Rollback is needed when a deployed term causes unexpected behaviour: a false-positive substitution in content that previously transcribed correctly, a confidence conflict with an adjacent entry, or a phonetic variant that produces a different incorrect output rather than the intended correct one. The rollback procedure is:

  1. Identify the term causing the issue from the QA error log or from a specific complaint. Confirm that the issue began after a specific version release (version log confirms when the term was deployed).
  2. Notify the Vendor Bridge that the specific term should be reverted to its pre-release state. The Vendor Bridge deploys the revert as a patch version.
  3. Any content captioned between the release and the revert that shows the false-positive pattern should be identified from the job log (all jobs with the affected glossary version in the date range) and flagged for QA review. Re-delivery decisions are made based on whether the false positive is present in the specific content type affected.
  4. The problematic term is returned to the approval chain with a note on the observed failure mode, so the phonetic variant or confidence score can be corrected before re-deployment.

The rollback procedure is not a sign of poor maintenance; it is a sign of a programme that has sufficient monitoring to detect when a deployed term causes regression. Programmes without version control and per-job logging cannot perform rollbacks because they cannot determine which jobs are affected, which makes a local failure into a pervasive one.

LMS catalog synchronisation

The caption glossary and the LMS course catalog are connected more tightly than most caption programmes recognise. Every course in the catalog represents a vocabulary profile — the specific set of terms that are relevant to that course’s content. When the catalog changes (courses added, retired, renamed, or migrated), the vocabulary implications for the glossary should be assessed. Without this connection, the glossary accumulates orphaned entries that serve no live content, and active content may lack glossary coverage because its vocabulary was never onboarded when the course was created.

Triggering a glossary review on course archive

When a course is retired from the LMS catalog, the Term Custodian should receive an archive notification and run a brief impact assessment:

  1. Identify all glossary terms that were linked to the archived course in the submission log (the “linked course or module” field from the submission form records this at submission time).
  2. For each identified term, check whether any other active courses also use that term (query the submission log for other courses that submitted or benefit from the same term).
  3. Terms used exclusively by the archived course become candidates for deprecation. If the archive content scope is closed (the course will no longer be captioned or re-captioned), the deprecation date is the archive date and the sunset date can be set to the organisation’s content retention expiry.
  4. Terms shared with active courses remain active and need no change.

The critical mistake to avoid is automatically deprecating all terms linked to an archived course without checking for shared use. This is the mirror of the archive content problem described in the deprecation section: a term that seems course-specific may actually appear in three other active courses, and premature deprecation introduces a vocabulary gap in those courses.

Course renaming and glossary impact

When a course is renamed in the LMS catalog — a common event when products are rebranded, when module titles are updated to reflect new regulatory requirements, or when a training series is restructured — the captioning implications depend on whether the course’s audio content is being re-recorded or only the metadata (title, description, catalog entry) is changing.

For metadata-only renames: no glossary action required. The audio and captions are unchanged. The LMS catalog ID used in the job log should be updated to reflect the new course name for future reference, but no vocabulary update is triggered.

For courses being re-recorded with updated content: the standard new-content vocabulary review applies. The Term Custodian reviews the updated content outline or script for new vocabulary that requires glossary additions, and the submission workflow is initiated before the recording session, not after.

LMS migration events

An LMS migration — moving from one LMS to another, or consolidating two LMS platforms after an acquisition — creates a concentrated vocabulary review opportunity. The migration event touches every course in the catalog simultaneously, which means it is an efficient moment to audit the entire glossary against the catalog that will be active in the new LMS.

The LMS migration vocabulary audit has three components:

The catalog-version lock

The concept of a catalog-version lock is the LMS-catalog equivalent of per-job version logging. When a course version is locked in the LMS (marked as final, submitted for regulatory review, or published to learners), the glossary version active at that lock date should be recorded in the course’s metadata. This creates a stable reference point: if the course is later re-captioned (for example, if the platform changes or the caption format changes), the catalog-version lock identifies which glossary version should be used for re-captioning to maintain consistency with the original production.

Without the catalog-version lock, a re-captioning request on a 3-year-old course uses the current glossary version, which may have deprecated terms that were correct at the time of the original production and added terms that create false positives in the content. The catalog-version lock prevents this by creating an explicit instruction: “re-caption this course using glossary v3.2.1, the version active at original production, not the current version.”

Quarterly audit cadence

The quarterly audit is the mechanism that prevents reactive maintenance from becoming the only kind of maintenance. A glossary maintained only in response to complaints, re-delivery requests, or obvious accuracy failures is a glossary that is always one quarter behind the vocabulary state the organisation needs. The quarterly audit proactively surfaces the vocabulary gaps, orphaned entries, and drift patterns that reactive maintenance misses.

The audit has four components:

Component 1: Accuracy-drift detection

Pull the QA accuracy scores from the past quarter (the monthly tracking described in the caption quality error rate calculator post) and segment them by content type. Identify content types where accuracy has declined relative to the prior quarter. A decline in accuracy on a specific content type without a corresponding decline on other types is often a glossary gap signal: new vocabulary has entered that content type that is not in the glossary. Compare the error log entries for the declining content type against the current glossary to identify which terms are generating substitution or deletion errors.

Component 2: Unused term review

Review the glossary for terms that have not appeared in any QA error log entry as a correction candidate in the past two quarters. A term that is never invoked may be: correct and working as expected (the caption model handles it without needing a glossary correction); no longer in use in current content (a deprecated technology or discontinued feature); incorrectly entered (the phonetic variant doesn’t match how the term is actually spoken in the organisation’s content). Unused terms with high confidence scores are particularly worth investigating — if the model hasn’t needed to apply them, the high confidence score may be creating false positives on adjacent vocabulary. Unused terms should be flagged for Domain Authority review rather than deleted immediately, since the apparent non-use may be because content where the term appears hasn’t been in the QA sample set.

Component 3: Missing term identification

Review the QA error logs from the past quarter for substitution errors that are not explained by known glossary gaps. These uncategorised substitution errors are the strongest signal of vocabulary that should be in the glossary but isn’t. Group them by content type and submit the most frequently occurring ones through the standard urgent-to-standard workflow depending on frequency and content sensitivity. The caption feedback loop post covers how to close the loop from QA error identification to glossary update to accuracy improvement in a structured way; the quarterly audit formalises the missing-term identification step in that loop.

Component 4: Sunset date review

Review all deprecated entries for entries within 3 months of their sunset date (as described in the deprecation section). For each:

Quarterly audit outputs

The quarterly audit produces three outputs that should be retained in the programme records:

Output Contents Audience
Audit summary report Terms added (count and type), terms deprecated, terms retired, unused terms reviewed, missing terms identified, accuracy trend by content type Term Custodian, L&D lead, programme records
Action log Specific terms submitted, Domain Authority review outcomes, Vendor Bridge deployments triggered, version releases made Term Custodian, Vendor Bridge
Annual review input Aggregated quarterly audit data formatted as glossary health section for the annual programme review Annual review participants (see annual review post)

The quarterly audit should be blocked as a recurring calendar event for the Term Custodian: 2–3 hours per quarter for most programmes. Programmes with large catalogs (100+ active courses) or high content velocity (20+ new modules per month) should budget 4–6 hours. The audit cannot be compressed below a threshold — the missing-term identification step requires reviewing error logs carefully, which takes time proportional to production volume. Organisations that try to run a quarterly audit in 30 minutes are running a checklist, not an audit.

Eight failure modes in glossary maintenance

These are the most common ways glossary maintenance programmes fail to deliver their intended accuracy benefit.

  1. No assigned ownership: maintenance defaults to “whoever notices a problem”

    When no individual is accountable for glossary maintenance, it becomes entirely reactive. Problems are fixed when someone complains. New terms are added when a producer escalates. Deprecation never happens because no one has the time to do it and it doesn’t feel urgent. The accuracy improvement from the initial glossary build erodes within two quarters. Assigning a named Term Custodian with a time allocation for maintenance (even 2–4 hours per week in a medium-size L&D team) is the single most effective structural change a programme can make after initial setup.

  2. Product team does not notify the Term Custodian before announcement

    Without a standing connection between product marketing and the glossary workflow, rebranding notifications arrive after the announcement — or not at all. The first discovery that a product was renamed is often a QA reviewer catching a substitution error in a post-announcement training module. Fixing this requires adding the Term Custodian to the product launch checklist as a required step before announcement, not after. This is a governance policy fix, not a maintenance workflow fix — it has to be codified in the captioning programme policy that product marketing has signed off on.

  3. Deprecated terms deleted instead of flagged as deprecated

    Deleting a deprecated term immediately fixes new content at the cost of breaking archive content. This is the most common single failure mode in rebranding maintenance. The symptom is a cluster of QA errors on older modules shortly after a rebranding, where the captions now contain incorrect transcriptions of the old product name because the ASR model has no recognition path for it. The fix requires restoring the deprecated entry (from version history, if version control was in use) and applying the correct deprecation-without-deletion process going forward. Programmes without version control cannot restore the deleted entry easily and must re-caption affected archive content from scratch.

  4. Acquisition vocabulary onboarded on a single-priority, all-at-once basis

    Trying to onboard all acquired vocabulary at once creates a large, unverified batch that the Vendor Bridge must deploy before it has been properly reviewed for conflicts and phonetic accuracy. The typical outcome is a deployment that fixes some terms, creates false positives from others, and then requires a series of patch rollbacks that introduce their own inconsistencies. The tiered priority model described above prevents this by sequencing the onboarding so that the highest-impact terms are deployed first, tested, and confirmed before the next tier is deployed.

  5. Version control implemented at the glossary level but not at the per-job level

    A glossary with version history but without per-job logging can answer the question “what was in version 4.2.1?” but not the question “which version was used to caption course HR-2024-0031?” The DCMP audit trail requirement is answered by the per-job log, not the version history alone. Implementing per-job logging is a one-time operational change: add a version field to the caption job submission form or to the job creation API call, and capture it in the job record at submission time. Retrofitting per-job logging is much harder than building it in from the start.

  6. LMS archive events not connected to glossary review

    When courses are retired from the LMS without a corresponding glossary review, the glossary accumulates orphaned entries at a rate that matches content retirement velocity. After 2–3 years, programmes without this connection have a glossary that is 30–50% orphaned entries for retired content — entries that take up confidence score bandwidth, create potential false-positive conflicts with active vocabulary, and make the quarterly audit harder to interpret. Establishing a notification or checklist step in the LMS content retirement process that prompts the Term Custodian to run the orphan assessment costs one line in the retirement procedure.

  7. Quarterly audit deferred to “when there’s time”

    The quarterly audit is the only maintenance activity that proactively identifies vocabulary gaps before they produce caption errors. When the audit is deferred, the programme operates in purely reactive mode: problems surface through QA failures and learner feedback rather than through structured review. A quarterly audit deferred for two quarters becomes a 6-month backlog that takes twice as long to clear as two separate audits would have taken, because the missing terms, unused terms, and orphaned entries have compounded over the longer period. Blocking the audit as a fixed calendar event with a formal time allocation prevents this compounding.

  8. Glossary shared with the captioning vendor without access controls or change management

    When the captioning vendor has direct write access to the production glossary — a common configuration when the glossary is hosted in the vendor’s platform — there is a risk that the vendor makes changes to improve their own accuracy metrics on vendor-evaluation content without documenting the changes or running them through the Term Custodian approval workflow. The fix is to require that all glossary changes, including changes made by the vendor, are routed through the Term Custodian for review before deployment, and that the version log reflects who made each change. The vendor SLA checklist post covers the contractual language that makes this a vendor obligation rather than a request.

FAQ

How large does a glossary typically get before the quarterly audit becomes unmanageable?
A quarterly audit on a glossary of 500 active terms takes 2–3 hours for a trained Term Custodian. At 1,000 active terms, 4–6 hours. Above 1,500 active terms, the audit should be split into domain-specific sub-audits (one per content vertical) rather than conducted as a single review. Programmes that reach 2,000+ active terms are typically large enterprises with 300+ courses; at that scale the quarterly audit is a team function rather than a single-person task. The most effective way to keep glossary size manageable is rigorous context restriction: a term that applies only to medical content should never appear in the general or technical content domains. Loose context restrictions cause all-domain inflation that multiplies the audit workload without improving accuracy on the specific content that needs the term.
Should the glossary be stored internally or hosted by the captioning vendor?
The authoritative copy of the glossary should be held internally, regardless of where the operational copy that the caption model reads is hosted. The internal copy is the source of record for the version log, the deprecation history, the per-job audit trail, and the rollback procedures. If the captioning vendor hosts the operational copy (which is common, since the model reads from the vendor’s glossary store), the Term Custodian should export a full copy of the glossary after every version release and store it in an internal location under version control. This export prevents single-vendor dependency: if the vendor relationship ends, the organisation retains the full glossary history and can onboard a new vendor without starting from scratch. The glossary-biased captioning post covers the technical architecture options for glossary storage across different caption service models.
What is the relationship between the caption glossary and the organisation’s company style guide or brand vocabulary guide?
They serve different purposes and should be maintained separately. The brand vocabulary guide is a reference document that tells writers and marketers how to refer to products and features consistently in written content. The caption glossary is an operational input to the ASR model that governs how audio is transcribed. The brand vocabulary guide is the source of record for what a term should look like; the caption glossary is the mechanism for making sure the ASR model produces that form. In practice, the best workflow is to subscribe the Term Custodian to the brand vocabulary guide update feed (or to include the Term Custodian in the distribution of brand guide change notifications) so that every brand vocabulary change is automatically evaluated for whether it has a glossary implication. Not all brand vocabulary changes have glossary implications — a change in how to describe a category of products in marketing copy does not necessarily change the spoken product names in training video — but the evaluation step takes 5 minutes and prevents the gap that arises when brand vocabulary and caption vocabulary drift apart.
How should glossary terms be handled for translated caption tracks?
Product names, proper nouns, acronyms, and SDK/API names typically do not translate — they appear in the translated caption track in their original English form (or the official localised form, if the organisation has one). These terms should be added to the translated-language caption model using the same approval workflow as English terms, but with the target-language phonetic variants rather than English phonetic variants. Regulatory vocabulary and clinical terminology may have official translated equivalents that differ from direct translation; these should be confirmed with the Domain Authority who covers the relevant jurisdiction or clinical specialty. The multilingual caption workflow post covers the full translation pipeline; the glossary maintenance workflow for translated tracks follows the same ownership and version-control model as English, with language-specific Domain Authorities for each target language.
What happens when an acquired company had a better glossary architecture than the acquiring company?
This situation arises in acquisitions where the acquired company had a more mature caption programme — it is more common than might be expected when a smaller company with an L&D focus is acquired by a larger company that built its captioning programme more recently. The correct approach is to document both glossary architectures and conduct a structured merge rather than defaulting to the acquiring company’s architecture. The merge assessment should compare: confidence scoring models, context restriction conventions, phonetic variant format standards, version control practices, and the quality of the deprecated term archive. In some cases, adopting the acquired company’s architecture for specific domains (particularly if the acquired company has superior coverage in its specialist vertical) produces better outcomes than forcing the acquired vocabulary into a less mature framework. This is a Term Custodian and Domain Authority decision, not a vendor decision.
How does the version log connect to an OCR accessibility complaint investigation?
An OCR (Office for Civil Rights) complaint investigation typically asks the responding organisation to demonstrate that it provides accessible content to individuals with disabilities — specifically, that caption tracks meet WCAG 2.1 AA accuracy standards. The version log supports this response in two ways: it demonstrates that the caption programme used a domain-specific glossary (not just default ASR) for the content in question, which is evidence of reasonable effort to achieve accuracy; and it allows the organisation to identify the exact glossary state at the time of production, which can be compared against the current glossary to show whether the accuracy failure (if any) resulted from a vocabulary gap that has since been corrected. Organisations that can show a documented version history, a quarterly audit cadence, and a per-job log are in a materially stronger compliance posture than organisations that can only show that captions exist. The building a caption compliance programme post covers the full OCR response framework.
Is there a minimum viable glossary maintenance process for a team of two that produces 10 modules per month?
For a two-person L&D team at 10 modules per month, the minimum viable process is: one person designated as Term Custodian (the accessibility-adjacent one, even if informally); a shared intake document (a simple spreadsheet or form-based doc) for term submissions; a monthly batch review rather than weekly (at 10 modules per month, the vocabulary change rate is low enough to batch monthly without causing accumulation problems); a per-job log field added to the caption job submission checklist (just one column: “glossary version”); and a semi-annual audit rather than quarterly. The full framework described in this post is sized for a 50–200-person L&D organisation. A two-person team should implement the principles (ownership, submission workflow, deprecation-not-deletion, per-job logging) without the procedural overhead designed for larger teams. The single most important element at any scale is the pre-announcement rebranding workflow: even a two-person team producing 10 modules per month needs to know about product rebrandings before the announcement, not after.

Glossary maintenance built into every caption job

GlossCap manages the term submission queue, version tracking, per-job glossary logging, and quarterly audit workflow — so your Term Custodian runs the programme rather than the spreadsheets. Every caption job records the glossary version it used. Every product rebranding flows through a structured update that preserves your archive content and prevents announce-day caption errors.

See GlossCap pricing Try the live demo

Other tools from the same factory: