Perform detailed technical analysis using streaming manifests, player logs, encoder metrics, QoS/QoE analytics, SCTE-35 signaling, transport stream analysis, and cloud monitoring tools.
Coordinate technical response across Engineering, Platform, Network Operations, CDN providers, cloud vendors, and third-party partners to restore service quickly.
Lead technical bridge calls during critical incidents, providing real-time updates, impact assessments, mitigation strategies, and recovery coordination.
Support Live Event Operators by providing technical guidance during event execution, validating mitigation actions, and assisting with operational decision-making.
Perform pre-event technical validation, redundancy verification, workflow testing, and production readiness reviews.
Conduct post-event root cause analysis (RCA), identify systemic issues, and drive corrective actions to improve platform reliability.
Develop and maintain troubleshooting documentation, operational runbooks, monitoring procedures, and knowledge base articles.
Support certification, onboarding, and production rollout of new streaming technologies, cloud workflows, monitoring tools, and vendor integrations.
Participate in on-call rotations supporting global live streaming operations and high-severity production incidents.
Requirements:
5+ years supporting large-scale live streaming, OTT platforms, broadcast technology, or video delivery systems
Experience serving as a technical escalation resource supporting live operations teams during Tier-1 live events
Strong troubleshooting experience across end-to-end streaming architectures
Deep understanding of HLS, DASH, CMAF/fMP4, MPEG-TS, SCTE-35, timed metadata, and adaptive bitrate streaming
Experience analyzing manifests, transport streams, encoder logs, player telemetry, QoS/QoE metrics, and cloud infrastructure metrics
Familiarity with AWS Media Services, CDN technologies, cloud networking, and streaming observability platforms
Working knowledge of contribution protocols including SRT, Zixi, RIST, RTP, UDP, and multicast workflows
Excellent incident management, root cause analysis, and cross-functional communication skills in high-pressure live production environments.