โšก New

Senior Manager Engineering

Albertsons Companies India

BengaluruFull-timeMid LevelOn-site

Job Description

Senior Engineering Manager, AIOps Platform

Role Overview:

We are seeking a visionary Senior Engineering Manager to lead our AIOps Platform Engineering organization. This executive-level role is responsible for the strategy, architecture, and operational excellence of our enterprise observability and AI-driven operations platform. You will own the end-to-end platform that powers multi-tenant data pipelines, enables predictive intelligence, and drives our Enterprise Service-Oriented Business (ESOB) strategy.

Leading a high-caliber team of platform engineers, software engineers and ML engineers, you will be the accountable executive for platform availability, cost-to-serve, and the AI-first transformation of how we observe, detect, and remediate at scale.

Key Responsibilities:

  • Define and execute the technical vision for the enterprise AIOps and observability platform, ensuring architectural runway for multi-tenant data ingestion, real-time analytics, and AI/ML inference at scale.
  • Own the 99.99% availability mandate across all platform services. Establish and govern SLOs, SLIs, and error budgets. Drive blameless postmortem culture and systemic reliability improvements.
  • Lead the development and operationalization of anomaly detection, event correlation, and automated remediation capabilities. Ensure model observability, explainability, and trust in production AI systems.
  • Architect and govern high-throughput, low-latency data pipelines supporting petabyte-scale telemetry ingestion. Ensure tenant isolation, data governance, and cost-efficient storage/compute.
  • Advance GitHub Actions-based CI/CD maturity, enforcing secure-by-default deployment pipelines, automated quality gates, and progressive delivery mechanisms.
  • Define and enforce observability standards across the technology portfolio, leveraging Grafana as the primary visualization and alerting backbone integrated with the enterprise toolchain ecosystem.
  • Build, mentor, and retain a world-class engineering organization. Define career architectures for Platform Engineers and SREs. Drive workforce planning and succession strategy.
  • Partner with Product Management, Data Science, and Business Unit leaders to align platform investments with enterprise priorities. Present technical health, risk posture, and strategic roadmaps to the executive committee and board.
  • Own the cost-to-serve model for observability infrastructure. Drive capacity planning, FinOps discipline, and vendor consolidation to maximize return on platform investment.
  • Ensure production readiness standards, disaster recovery/business continuity planning, and adherence to regulatory requirements for data retention, privacy, and auditability.

Required Qualifications:

  • 12+ years of progressive technology experience, including 5+ years in engineering leadership roles managing platform, infrastructure, or SRE organizations within large-scale enterprise environments.
  • Demonstrated deep expertise in Site Reliability Engineering principles, including error-budgeting, chaos engineering, incident command, and reliability-centered design. Prior hands-on SRE experience is explicitly required.
  • Proven track record building or scaling enterprise observability platforms (metrics, logs, traces) and operationalizing AIOps capabilities such as anomaly detection, root-cause analysis, and intelligent alerting.
  • Expert-level proficiency with Grafana (Grafana Cloud or Grafana Enterprise), including dashboard design, alerting configuration, data source integration, and performance optimization at enterprise scale.
  • Extensive experience architecting multi-tenant, high-availability data pipelines using modern streaming and batch technologies.
  • Deep familiarity with GitHub Actions and modern DevOps practices, including infrastructure-as-code, automated testing strategies, and progressive deployment patterns.
  • Experience building and operating ML feature stores, model serving infrastructure, and observability stacks for production machine learning systems.
  • Exceptional written and verbal communication skills with demonstrated ability to present to C-suite executives, board members, and cross-functional leadership teams.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field; advanced degree preferred.

Preferred Qualifications:

  • Experience driving ESOB (Enterprise Service-Oriented Business) or similar enterprise-wide service transformation strategies.
  • Background in large-scale distributed systems operating at 99.99% availability or higher in regulated industries (financial services, healthcare, telecommunications).
  • Contributions to open-source observability or AIOps communities.
  • Familiarity with OpenTelemetry, Prometheus, Loki, Tempo, and cloud-native observability architectures.
  • Experience with FinOps frameworks and cloud cost optimization at scale.
  • Prior accountability for vendor management and enterprise contract negotiation with observability and infrastructure providers.
  • Certifications such as AWS/Azure/GCP Solutions Architect, Certified Kubernetes Administrator (CKA), or SRE-specific credentials.

Leadership Competencies:

  • Ability to synthesize complex technical, business, and organizational dynamics into coherent platform strategies that create durable competitive advantage.
  • Commands respect in boardroom and executive committee settings. Influences without authority and builds consensus across matrixed organizations.
  • Recognized as a leader who attracts top-tier engineering talent, develops high-potential individuals, and builds resilient, autonomous teams.
  • Unwavering commitment to measurement, accountability, and continuous improvement. Sets the gold standard for operational excellence.
  • Drives large-scale organizational and technical change with clarity, empathy, and velocity. Navigates ambiguity and resistance to deliver outcomes.
  • Internalizes the voice of the enterprise customer and ensures platform decisions translate into measurable business value and user trust.
  • Champions the responsible adoption of artificial intelligence and machine learning to transform operational workflows, while maintaining rigorous governance and ethical guardrails.

Posted Today

Related Jobs

Related Searches

Apply Now