Site Reliability Engineer, Infrastructure Engineering
coreweaveu
Job Description
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability.
Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com .
We're proud to be a Living Wage accredited Employer.
What You'll Do
The MetalDev team within CoreWeave's Hardware Compute organisation develops software automation tooling and services used to bring up data centre rack systems and manage bare‑metal infrastructure. We provide core reliability, availability, and operational stability functions across regional data centres to ensure seamless infrastructure provisioning.
About the role
As a Site Reliability Engineer on the MetalDev team, you will split your focus between production operations and reliability (60%) and engineering automation (40%). You will lead incident response, troubleshooting, root‑cause analyses, and post‑incident reviews while participating in an on‑call rotation. In this role, you will write resilient Go code, build Prometheus and Grafana dashboards, and develop automated remediation workflows to reduce manual overhead.
Additionally, you will define SLOs and KPIs, improve CI/CD deployment pipelines, and create self‑service tooling for Fleet Operations and Hardware engineering teams.
Who You Are
Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience). 3+ years of experience in Site Reliability Engineering, production engineering, cloud infrastructure, or software engineering. Working proficiency in Go with experience developing production-quality software. Hands‑on production experience with Kubernetes and containerised microservices.
Experience with observability and telemetry stacks, specifically Prometheus and Grafana. Demonstrated track record supporting production services, leading incident management, and participating in on‑call rotations. Strong troubleshooting, analytical, and technical documentation skills.
- Experience managing or automating bare‑metal infrastructure.
- Familiarity with BMCs, Redfish, or server‑management technologies.
- Experience building automated remediation or self‑healing systems.
- Familiarity with public cloud platforms such as AWS or GCP.
Wondering if you're a good fit?
- You love to: Build automated remediation workflows and eliminate operational toil across bare‑metal infrastructure.
- You're curious about: Developing low‑latency telemetry pipelines and scaling self‑healing systems across massive data centre clusters.
- You're an expert in: Production incident management, Go programming, and designing robust Kubernetes‑native observability tools.
Why CoreWeave?
At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper‑growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values:
- Be Curious at Your Core
- Act Like an Owner
- Empower Employees
- Deliver Best‑in‑Class Client Experiences
- Achieve More Together
We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for take‑off, the organisation's growth opportunities are constantly expanding.
You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!
We're hiring across multiple levels. Typical cash compensation ranges from ~223,000-298,000 PLN , with additional performance based bonus & equity that can significantly increase total compensation. The starting salary will be determined by job-related knowledge, skills, experience, and the market location.
We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).
To fulfill our obligation to protect client data, successful applicants offered employment with CoreWeave will be required to complete a basic criminal record check, conducted in compliance with GDPR. Employment offers are conditional upon receiving satisfactory check results.
What We Offer
In addition to a competitive salary, we offer a variety of benefits
#J-18808-Ljbffr