--- name: migrate-a-live-production-platform category: media description: Migrate a live production platform while preserving signal, control, accessibility, rights, monetization, monitoring, and event recovery through parallel validation. Use when active broadcasts or streams cannot tolerate a big-bang tooling change. --- # migrate-a-live-production-platform Migrate complete live outcomes, not isolated features. ## When to use - Use for switcher, contribution, encoder, cloud production, graphics, playout, caption, distribution, player, or control-platform migrations. - Use ordinary equipment replacement when no multi-system control, live event, customer, or recovery contract changes. ## Preconditions - Establish production, broadcast engineering, network, venue, platform, accessibility, security, rights, ad, analytics, support, vendor, finance, and event authority. - Preserve signal diagrams, configurations, macros, scenes, graphics, keys, accounts, schedules, runbooks, recordings, captions, player settings, incidents, contracts, and quality baselines. - Define event calendar, protected shows, acceptable interruption, degraded mode, capacity, release waves, change freezes, rollback owner, and recovery communication. ## Procedure 1. Build a **workflow and dependency inventory** from source, contribution, sync, audio, switch, graphics, replay, captions, encoding, network, distribution, ads, DRM, player, chat, analytics, recording, archive, and support. 2. Record operators, permissions, devices, sites, accounts, credentials, rate limits, protocols, formats, timing, latency, regions, vendor dependencies, and common-mode failures. 3. Define **output and control compatibility** for resolution, frame rate, color, audio, loudness, captions, timecode, metadata, SCTE or ad markers, DRM, latency, playback, APIs, automation, audit, and operator actions. 4. Map every configuration, scene, macro, graphic, schedule, endpoint, key, access role, alert, and runbook to the target or explicit retirement. 5. Build the target in an isolated environment with scoped credentials, production-shaped sources, load, monitoring, logging, and recording. 6. Train operators on normal, degraded, failure, takeover, and rollback behavior; measure workload and prevent ambiguous dual control. 7. Run **parallel validation and cutover** through synthetic, rehearsal, shadow, internal, low-risk show, venue, region, and audience waves. Give each event a fenced, monotonically increasing publication epoch so manifests, ads, DRM policy, notifications, chat state, and canonical recordings reject writes from an older authority. 8. Compare signal quality, timing, audio, caption accuracy and delay, graphics, ads, rights, security, latency, player behavior, analytics, control response, and archive. 9. Rehearse source loss, network loss, encoder crash, platform region loss, caption failure, graphics error, account lockout, compromised key, operator failure, and vendor outage. 10. Define event-by-event go, no-go, cutback, recovery, communications, recording, and post-event reconciliation authority. 11. Execute cutover outside protected events where possible, verify every destination independently, and keep the prior path warm for the approved window. 12. Rehearse **rollback and event recovery** without losing recordings, captions, schedules, ad reconciliation, chat moderation, or accepted configuration changes. 13. Monitor live and post-event signal, audience, accessibility, security, rights, monetization, support, operator, capacity, and archive outcomes. 14. Retire the old platform only after protected events, configs, recordings, credentials, contracts, recovery, and operator proficiency have verified dispositions. ## Failure plan - If outputs disagree or both platforms can publish, fence one publication authority before continuing. - If rollback would break an event already started, keep the authoritative live path and recover supporting control planes around it. - If accessibility, rights, security, or recording fails, use the approved degraded mode or stop the affected output. - If vendor capacity or regional behavior is unproven, hold that audience or venue on the verified path. ## Worked example A live events network moves from an on-premises and legacy-cloud stack to a new production platform across venues and regions. The workflow includes remote contribution, switchers, graphics, captions, ads, DRM, chat, analytics, players, and archives. The team maps every control and format, builds isolated target capacity, validates shows in parallel without double publishing, trains operators through failure drills, canaries low-risk events, and rehearses event recovery that preserves recordings, captions, schedules, and monetization records. ## Done - A platform migration register records workflows, dependencies, owners, configurations, operators, formats, controls, events, waves, exceptions, and retirement gates - A signal, accessibility, and control parity report proves video, audio, timing, captions, graphics, rights, ads, security, player, analytics, automation, recording, and operator outcomes - A cutover and event recovery rehearsal verifies parallel validation, publication authority, failures, degraded modes, communications, recordings, reconciliation, rollback, and protected-event recovery