โšก New

Observability Engineer

Apptoza Inc.

TorontoFull-timeMid LevelOn-site

Job Description

Role: Observability Engineer (KUBERNETES & CONTAINERS, GITOPS & AUTOMATION, QUERY LANGUAGES, CLOUD & INFRASTRUCTURE)

Location: Toronto, ON โ€“ 4 days onsite

Contract Role

Job Description

ABOUT THE ROLE

We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. You'll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications.

This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure.

OBSERVABILITY STACK OWNERSHIP

  • Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents
  • Manage observability deployments using GitOps principles and infrastructure as code
  • Implement long-term metrics storage solutions with cloud object storage
  • Maintain and upgrade observability components across development, QA, UAT, production, and DR environments
  • Configure distributed observability architecture spanning multiple data centers and cloud providers

METRICS & MONITORING

  • Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications
  • Create ServiceMonitors, PodMonitors for automated metrics collection
  • Develop rules for intelligent alerting with minimal false positives
  • Configure multi-cluster metrics federation and aggregation
  • Optimize metrics cardinality, storage efficiency, and query performance
  • Implement recording rules for pre-aggregated metrics and SLI calculations

DASHBOARDS & VISUALIZATION

  • Build comprehensive Grafana dashboards for infrastructure health, application performance, and business metrics
  • Create reusable dashboard templates and libraries for development teams
  • Implement dashboard-as-code
  • Configure multiple datasources
  • Design executive dashboards with SLO/SLI tracking and business KPIs
  • Implement role-based access control and multi-tenancy in Grafana

LOGGING INFRASTRUCTURE

  • Deploy and manage centralized logging solutions (Loki, ELK, Splunk, or Similar)
  • Configure log collection agents (Promtail, Fluentd, FluentBit, Vector, etc.)
  • Design log retention policies balancing cost, compliance, and operational needs
  • Create LogQL/Lucene queries and log-based alerts
  • Implement log correlation with metrics and traces for unified troubleshooting
  • Build log aggregation pipelines with parsing, filtering, and enrichment

ALERTING & INCIDENT MANAGEMENT

  • Configure intelligent alerting with Alertmanager or equivalent platforms
  • Design alert rules with appropriate severity levels, thresholds, and SLOs
  • Implement alert routing to Slack, PagerDuty, ServiceNow, email, and webhooks
  • Create automated runbooks and remediation workflows
  • Develop alert inhibition, silencing, and grouping strategies
  • Tune alerting to achieve signal-to-noise ratio improvements
  • Integrate with incident management and on-call rotation systems

DISTRIBUTED TRACING

  • Implement distributed tracing solutions using OpenTelemetry, Jaeger, or Tempo
  • Instrument applications for trace collection and correlation
  • Configure trace sampling strategies for cost and performance optimization
  • Build trace-based dashboards for latency analysis and dependency mapping
  • Integrate tracing with metrics and logs for comprehensive observability

AI/ML FOR OBSERVABILITY

  • Implement AI-powered anomaly detection for metrics and logs
  • Build predictive alerting using machine learning models to forecast issues before they occur
  • Develop intelligent alert correlation and root cause analysis systems
  • Integrate LLM-based tools for log analysis and troubleshooting assistance
  • Implement AIOps capabilities for automated incident triage and resolution
  • Use AI to optimize alert thresholds and reduce false positives
  • Build natural language query interfaces for observability data
  • Implement Model Context Protocol (MCP) for AI agent integration with observability platforms

AUTOMATION & PLATFORM INTEGRATION

  • Automate observability deployment using GitOps workflows (FluxCD, ArgoCD)
  • Integrate with CI/CD pipelines for automated testing and validation
  • Build self-service portals for teams to create dashboards and alerts
  • Develop APIs and CLIs for observability automation
  • Integrate with secrets management solutions (Vault, AWS Secrets Manager)
  • Configure LDAP/AD/SSO authentication for observability platforms
  • Automate compliance reporting and audit logging

ENABLEMENT & COLLABORATION

  • Onboard application teams to observability platforms
  • Provide guidance on instrumentation best practices
  • Create documentation, training materials, and self-service guides
  • Conduct workshops
  • Support development teams during incidents with observability insights
  • Collaborate with SRE, DevOps, and platform engineering teams

PERFORMANCE & COST OPTIMIZATION

  • Monitor and optimize observability stack resource consumption
  • Implement autoscaling for stateless observability components
  • Tune data retention, compaction, and downsampling strategies
  • Conduct capacity planning for metrics and log storage growth
  • Optimize query performance and dashboard response times
  • Implement cost allocation and chargeback for multi-tenant environments

REQUIRED QUALIFICATIONS

EXPERIENCE

  • 4-6 years of experience in observability, monitoring, SRE, or platform engineering roles
  • 3+ years hands-on production experience with Prometheus and Grafana
  • 2+ years working with Kubernetes and containerized environments
  • Strong expertise with PromQL for metrics querying and alerting
  • Experience deploying and managing observability stacks at scale (1000+ nodes)
  • Proven track record of reducing MTTR through effective observability

CORE TECHNICAL SKILLS

OBSERVABILITY PLATFORMS:

  • Prometheus (including Prometheus Operator)
  • Grafana (dashboards, alerting, plugins)
  • Thanos, Cortex, or Mimir for long-term storage
  • Alertmanager or equivalent alerting platforms

LOGGING:

  • Loki, Elasticsearch/ELK Stack, Splunk, or CloudWatch Logs
  • Log collection agents (Promtail, Fluentd, FluentBit, Vector)
  • LogQL, Lucene, or equivalent query languages

TRACING:

  • OpenTelemetry (OTEL Collector, instrumentation)
  • Jaeger, Zipkin, Tempo, or AWS X-Ray
  • Trace sampling and correlation strategies

KUBERNETES & CONTAINERS:

  • Kubernetes architecture and operations (1.24+)
  • Custom Resource Definitions (CRDs) and Operators
  • ServiceMonitor, PodMonitor, PrometheusRule resources
  • Kubernetes metrics (kube-state-metrics, node-exporter, cAdvisor)

GITOPS & AUTOMATION:

  • Basic SQL for data analysis

CLOUD & INFRASTRUCTURE:

  • AWS, Azure, or GCP cloud platforms
  • Multi-cloud and hybrid architectures
  • vSphere or on-premises virtualization (nice to have)

AI/ML & EMERGING TECHNOLOGIES

#J-18808-Ljbffr

Posted Today

Related Jobs

Related Searches

Apply Now