⚡ New

Software Developer 4

Stanford University

Menlo ParkFull-timeMid LevelOn-site

Job Description

Job Description

The Rubin U.S. Data Facility at SLAC seeks a senior Software Developer / Site Reliability Engineer to lead and execute complex software and infrastructure work supporting Rubin Observatory production services. This role will help operate, automate, and improve the systems that support Alert Production, Data Release Production, database-backed services, Kubernetes applications, Weka and object-storage services, data movement, and other USDF-hosted operational platforms.

Job Description

The Rubin U.S. Data Facility at SLAC seeks a senior Software Developer / Site Reliability Engineer to lead and execute complex software and infrastructure work supporting Rubin Observatory production services. This role will help operate, automate, and improve the systems that support Alert Production, Data Release Production, database-backed services, Kubernetes applications, Weka and object-storage services, data movement, and other USDF-hosted operational platforms.

This position combines senior software development, site reliability engineering, production operations, deployment engineering, and cross-team technical leadership. The successful candidate will work across Rubin Data Management, SQuaRE, Prompt Processing, DRP, database, storage, networking, and USDF infrastructure teams to ensure that critical services can be deployed, monitored, debugged, upgraded, recovered, and scaled safely.

The role is aligned with the Software Developer 4 level: leading and executing difficult or complex programming and analysis work, contributing broad technical responsibility across multiple functions, and interfacing with other complex systems and programs.

Key Responsibilities

  • Lead and execute complex software, automation, and operational engineering projects for Rubin production services at USDF, including Alert Production, DRP, data access, databases, and service infrastructure.
  • Support and improve Kubernetes-hosted services deployed through Rubin’s Phalanx/GitOps environment, including Helm configuration, secrets management, environment promotion, release coordination, and operational rollback.
  • Maintain and troubleshoot PostgreSQL-backed services, including backup and restore, performance, schema-change coordination, monitoring, and incident response.
  • Help operate and debug services that depend on Weka storage, S3 gateways, object storage access, shared filesystems, Butler repositories, batch processing, and high-throughput data movement.
  • Develop strategies, methods, tools, and procedures that improve reliability, automation, reproducibility, observability, and operational efficiency across USDF production services, consistent with Software Developer 4 expectations for process improvement and innovation.
  • Direct or drive all phases of selected technical projects, including requirements analysis, design, implementation, testing, deployment, documentation, and operational handoff.
  • Lead testing, debugging, change control, reporting, and documentation for major service and infrastructure changes.
  • Build and maintain monitoring, logging, alerting, dashboards, runbooks, and operational documentation using tools such as Prometheus, Grafana, Loki, and related systems.
  • Provide subject-matter expertise for complex technical problems involving Kubernetes, databases, storage, networking, deployment automation, and distributed production services.
  • Recognize opportunities for high-impact, long-term reliability improvements that span team or organizational boundaries, and recommend or implement actions to resolve them.
  • Mentor technical staff and collaborate with developers, operators, database engineers, storage engineers, and scientists to improve service reliability and operability.

Desired Qualifications And Experience

  • Bachelor’s degree and ten years of relevant experience, or a combination of education and relevant experience, consistent with the Software Developer 4 profile.
  • Demonstrated experience designing, developing, testing, deploying, operating, and maintaining complex applications or infrastructure services.
  • Strong experience with Kubernetes, Helm, GitOps, CI/CD, Linux systems, Python, shell scripting, and production service automation.
  • Strong understanding of relational databases, especially PostgreSQL, including operational monitoring, performance troubleshooting, backup/restore, and schema-change coordination.
  • Experience with large-scale storage or data systems, ideally including Weka, S3/object-storage interfaces, shared filesystems, high-throughput data transfer, or scientific data repositories.
  • Experience with observability systems such as Prometheus, Grafana, Loki, alerting systems, logs, metrics, and operational dashboards.
  • Strong understanding of software development life cycle, quality-control practices, change management, incident response, and operational documentation.
  • Exceptional written and oral communication skills for technical and non-technical audiences, including the ability to work across distributed teams and organizational boundaries.
  • Experience with Rubin Observatory software, Phalanx, Butler, Qserv, Kafka, Prompt Processing, DRP workflows, Weka, S3DF/USDF, or large-scale astronomical data processing would be especially valuable.

Given the nature of this position, SLAC is open to on-site, hybrid, and remote work options

Responsibilities

Core Duties

  • Lead strategic planning for complex projects with significant business, regulatory and/or technical challenges requiring subject matter expertise.
  • Direct and drive all phases of a project, including systems analysis, program design, development, and implementation.
  • Consult with clients to understand, conceptualize, design, implement, and deploy technology solutions for difficult and complex applications.
  • Develop strategies, methodologies, and procedures to innovate and drive processes to improve efficiency and innovation.
  • Conceptualize and perform complex programming and analysis which typically involves multi-project leadership and/or broad responsibility.
  • Lead testing, debugging, change control, reporting and documentation for major projects on large-scale software solutions.
  • Lead professional or project staff as necessary, setting goals, facilitating agile development practices, and providing technical mentorship, working on all phases of application development projects.
  • Develop and engage in long-term strategic planning, managing strategies that may include relationship development, communications, and compliance.
  • Provide subject matter expertise as an expert in many technical areas for school/unit and/or the university regarding the development applications, troubleshooting technical problems, and the implementation of new processes and/or procedures.
  • Recognize opportunities for high-impact, long-term changes that address strategic goals. Recommend actions and/or resolve complex issues that span organizational boundaries.
  • May supervise or mentor other staff.

Minimum Education And Experience

Bachelor's degree and ten years of relevant experience, or a combination of education and relevant experience.

Knowledge, Skills And Abilities

  • Ability to quickly learn and adapt to new technologies and programming tools.
  • Demonstrated experience in designing, developing, testing, and deploying applications.
  • Strong understanding of data design, architecture, relational databases, and data modeling.
  • Exceptional and effective written and oral communication skills to address a wide variety of audiences, both technical and non-technical.
  • Demonstrated project management ability to employ integration, scope time management, cost, quality, human resources, communications, risk, and procurement components.
  • Thorough understanding of all aspects of software development life cycle and quality control practices.
  • Ability to define and solve logical problems for highly technical applications.
  • Ability to select, adapt, and effectively use a variety of programming methods.
  • Demonstrated resilience, diplomacy, influence, relationship building, and critical thinking skills in a variety of situations.
  • Ability to recognize and recommend needed changes in user and/or operations procedures.

#J-18808-Ljbffr

Posted Today

Related Jobs

Related Searches

Apply Now