⚑ New

Manager

Tata Communications

Tamil NaduFull-timeMid LevelOn-site

Job Description

1 Year Contract role


Role: Kubernetes Manager

Level: L3 / Senior Kubernetes Administration & Operations


Primary Focus: Linux Administration, Kubernetes Administration, Production Operations, SRE, Monitoring, Security, Upgrades and Customer Escalation


Mandatory Skills

  • CKA certification (Preferred)
  • Strong/deep knowledge of Linux administration and troubleshooting
  • Strong hands-on experience in Kubernetes administration and L3 troubleshooting
  • Knowledge on virtualization like vSphere, OpenStack, Ovirt , KVM
  • Production experience in Kubernetes cluster operations, incident management and troubleshooting
  • Strong knowledge of Kubernetes architecture, control plane, worker nodes, networking, storage and workloads
  • Ability to handle P1/P2/P3 incidents and customer escalations within SLA
  • Hands-on experience with Kubernetes upgrades, patching and maintenance
  • Strong understanding of container runtime technologies:
  • containerd
  • CRI-O
  • Docker
  • CRI
  • Experience with container registries, preferably Harbor

Monitoring & Observability

  • Configure and administer Prometheus
  • Configure and manage Alert manager
  • Grafana dashboard creation and monitoring
  • Kubernetes cluster health monitoring
  • Alert investigation and resolution within SLA
  • Alert tuning and noise reduction
  • Identify duplicate/repeated alerts and optimize alert rules
  • Experience with ELK/EFK and Fluentd
  • Log collection, analysis and troubleshooting

CI/CD & DevOps

Hands-on experience with:

  • Jenkins
  • GitLab CI/CD
  • Git
  • Container image build and deployment
  • Kubernetes deployment automation
  • CI/CD troubleshooting

Kubernetes Operations

Responsible for:

  • Day-to-day Kubernetes cluster administration
  • L3 troubleshooting of Kubernetes issues
  • Kubernetes upgrade and patching activities
  • Cluster health checks
  • Node and pod troubleshooting
  • Control-plane and worker-node troubleshooting
  • Kubernetes networking troubleshooting
  • Kubernetes storage troubleshooting
  • Resource and performance analysis
  • Incident/problem management
  • Root Cause Analysis (RCA)
  • Vulnerability remediation
  • Production change implementation
  • Customer-specific Kubernetes requirements

Security & Vulnerability Management

  • Kubernetes security best practices
  • Vulnerability identification and remediation
  • Container image vulnerability management
  • Registry security
  • Kubernetes RBAC
  • Network/security policy concepts
  • SSL/TLS certificate management
  • Experience with security tools such as Gatekeeper
  • Ability to coordinate vulnerability fixes across OS, Kubernetes and container layers

Networking

Strong understanding of:

  • DNS
  • TCP/IP and L3 networking
  • Load Balancers
  • SSL/TLS termination
  • Ingress
  • Kubernetes Services
  • Network troubleshooting
  • Ingress Controllers
  • Istio Gateway
  • Service Mesh concepts
  • NGINX
  • Kong API Gateway

Storage

Good understanding of Kubernetes storage and underlying infrastructure:

  • NFS / File storage
  • Block storage
  • Object storage
  • PV/PVC
  • Storage Class
  • CSI
  • Storage troubleshooting
  • Mount and I/O issues
  • Storage performance concepts


Platform Components

Working knowledge of:

Component Required Knowledge


Istio

Service Mesh, Gateway, traffic management

NGINX

Ingress / reverse proxy

Kong

API Gateway

Redis

Cache / datastore

PostgreSQL

Database fundamentals & troubleshooting

Kafka

Messaging / event streaming

RabbitMQ

Message broker

Keycloak

IAM / authentication

Gatekeeper

Kubernetes policy enforcement

Harbor

Container registry

Prometheus

Monitoring

Grafana

Visualization

Alertmanager

Alert management

ELK

Centralized logging

Fluentd

Log collection

Customer & Operations Responsibilities

  • Provide 24Γ—7/on-call customer support as required
  • Handle customer escalations at OS and Kubernetes levels
  • Understand and resolve customer production issues
  • Participate in incident bridges and technical discussions
  • Perform RCA for critical incidents
  • Track and resolve tickets within defined SLA
  • Handle day-to-day customer requirements
  • Coordinate with development, compute, network, storage, security and application teams
  • Ensure production changes follow the approved change-management process
  • Prepare technical documentation and operational procedures

Posted Today

Related Jobs

Related Searches

Apply Now