⚑ New

L3 Support Engineer

Comply India

CochinFull-timeMid LevelOn-site

Job Description

Who We Are:

We areCOMPLY.

For compliance people.

We pride ourselves on being the champion for compliance professionals. We help clients navigate the ever-changing regulatory environment by merging technology, consulting, and education; We serve more than 7,000 clients globally through our solutions, including ComplySci, RIA in a Box, National Regulatory Service (NRS), and illumis. Our high-growth organization has been recognized with numerous awards, including Inc. 5000, Institutional Asset Manager Awards, Private Equity Wire Awards, and the Women in Data & Technology Awards.

Production Support Engineer:

Comply is building a world-class production support organization. We value engineers who take ownership, care deeply about reliability, and are driven by data. We're looking for pragmatic problem-solvers who thrive in fast-paced environments and are committed to continuous improvement.

About the Role:

We are seeking a skilled Production Support Engineer to join our rapidly growing platform team. In this role, you will take ownership of diagnosing and resolving complex production issues, troubleshooting application code using proven software development techniques, and helping drive down Mean Time to Resolution. You'll work collaboratively

with developers, infrastructure teams, and product leadership to maintain system reliability while identifying opportunities for continuous improvement. This is a hybrid role sitting at the intersection of operations and engineering you'll spend roughly equal time on incident response and hands-on engineering work, including code-level debugging, internal tooling development, and automation. This is an excellent opportunity to develop your expertise in production systems, build on your troubleshooting skills, and grow into a senior or leadership position within a high-performing team.

Responsibilities:

  • Production Issue Diagnosis & Resolution: Take ownership of incident triage and troubleshooting. Use software development techniquesdebugging, log analysis, profiling, code inspectionto diagnose root causes of application failures. Escalate complex issues appropriately while working to resolve issues within SLA targets. In some cases, implement targeted fixes, patches, or workarounds directly in the codebase in collaboration with development teams
  • Application Code Troubleshooting: Perform hands-on analysis of application code to identify bugs, performance bottlenecks, and misconfigurations. Use debuggers, APM tools, and systematic methodologies to isolate issues. Collaborate with development teams on complex problems that require code-level investigation.
  • Incident Response & Management: Respond to production alerts and incidents with urgency and professionalism. Follow established incident response procedures, communicate status clearly to stakeholders, and document resolution steps for knowledge sharing.
  • Post-Incident Review & Learning: Participate in blameless post-mortems following critical incidents. Contribute to root cause analysis efforts, identify systemic improvements, and help implement corrective actions to prevent recurrence.
  • Runbook & Documentation Development: Create and maintain comprehensive runbooks, playbooks, and troubleshooting guides. Document common issues, resolution steps, and architectural context to empower the team and accelerate future incident response.
  • Observability & Monitoring: Develop familiarity with our observability stack and monitoring systems. Alert on anomalies proactively and help identify gaps in monitoring coverage. Work with platform teams to improve visibility into system behaviour.
  • Automation & Efficiency: Identify repetitive manual tasks and contribute to automation efforts using scripting (Python, Bash, or similar). Build tools and utilities that reduce MTTR and eliminate operational toil.
  • Cross-Functional Collaboration: Work effectively with development, infrastructure, and product teams. Communicate technical findings clearly to both technical and non-technical stakeholders. Escalate issues appropriately and contribute to architectural discussions affecting production stability.
  • Continuous Learning: Develop your expertise in production operations, software engineering practices, and system architecture. Seek feedback, embrace mentorship, and actively pursue opportunities to expand your technical skills.

Qualifications:

Required Experience:

  • 3-5 years of hands-on production support, DevOps, site reliability engineering, with a demonstrable engineering component writing tools, automation, or application- level fixes, not solely reactive support.
  • Demonstrated ability to troubleshoot application issues using software development techniques: code analysis, debugging, log inspection, performance profiling, and systematic problem-solving.
  • Working knowledge of SaaS or cloud-native application architectures and operational challenges.
  • Experience responding to production incidents in fast-paced environments; understanding of incident severity levels and escalation procedures.
  • Strong diagnostic and analytical skills: ability to move methodically from symptom observation to root cause identification.
  • Proficiency with at least one programming or scripting language (Python, Go, Bash, Java, etc.).
  • Comfortable with Linux/Unix command-line environments and basic system administration.
  • Experience with code review practices, Git workflows, and contributing to engineering team repositories alongside product developers.

Advantageous:

  • Experience with containerized systems (Docker, Kubernetes) or cloud platforms (AWS, Azure, Google Cloud).
  • Familiarity with observability/APM tools (Datadog, New Relic, Splunk, Prometheus/Grafana, or similar).
  • Background in software development or QA automation (ability to read and understand application code eg.SQL).
  • Experience with incident management frameworks, blameless post-mortems, or RCA methodologies.
  • Knowledge of CI/CD pipelines, deployment processes, or infrastructure-as-code practices.
  • Track record of mentoring junior team members or contributing to knowledge-sharing initiatives.

Key Skills & Attributes:

  • Application Troubleshooting: Solid ability to diagnose failures through code review, log analysis, and tracing. Comfortable diving into unfamiliar codebases.
  • Technical Depth: Solid understanding of HTTP/networking, databases, caching, message queues, and common architectural patterns in modern applications.
  • Analytical Mindset: Methodical problem-solving approach; ability to form hypotheses and test them systematically rather than guessing.
  • Scripting & Automation: Comfort writing or maintaining scripts to automate diagnostics, data collection, or operational tasks.
  • Communication: Ability to explain technical issues clearly to non-technical stakeholders; document decisions and findings for team reference.
  • Ownership: Proactive in taking responsibility for issues, following them through to resolution, and seeking closure rather than deferring.
  • Collaboration: Works well with distributed teams; asks clarifying questions and doesn't hesitate to reach out for help when needed.
  • Continuous Learning: Actively develops technical skills, seeks feedback, and embraces learning from incident reviews and peer knowledge-sharing.

Growth & Development:

This role is designed as a growth opportunity. You will have the chance to:

  • Develop deep expertise in production troubleshooting, system architecture, and operational best practices.
  • Learn from incident reviews and mentorship from senior engineers and team leadership.
  • Take on expanded responsibilities as you mature, such as on-call rotation leadership, runbook ownership, or mentoring junior engineers.
  • Progress toward senior engineer or team lead positions within the support organization.

What We Offer:

At Comply, you'll join a team that values technical excellence and operational discipline.

We provide:

  • Exposure to complex, high-scale SaaS systems with real operational challenges.
  • Mentorship from experienced engineers and technical leaders.
  • Opportunities to influence production reliability and drive operational improvements.
  • A data-driven, metrics-focused culture that rewards continuous improvement.
  • Remote-friendly work environment with flexibility to suit your circumstances.

Posted Today

Related Jobs

Related Searches

Apply Now