⚡ New

Databricks Engineer

AdvanceWorks

RemoteFull-timeMid LevelRemote

Job Description

Challenge yourself and us

At AdvanceWorks, we learn together, grow together, and have fun together. And we do all this while creating great software solutions for clients across many industries and geographies.

We are constantly challenging ourselves and each other to use the latest technologies and the best methodologies and make sure we walk the talk.


What will be your role

We are looking for a Databricks Data Engineer to join our team in Lisbon and contribute to the development and evolution of modern data ingestion and Change Data Capture (CDC) solutions for a global organization operating at scale.

In this role, you’ll work on a modern data architecture built around Databricks, Apache Spark, Delta Lake, Azure, and Kafka, helping evolve a flexible and scalable approach to data ingestion across multiple source systems.

You’ll have an active role in defining and evolving ingestion patterns, developing Databricks notebooks and processing frameworks, configuring pipelines, and ensuring the integrity, traceability, and scalability of data throughout the process.

The project is moving from a traditional CDC approach focused primarily on Mainframe environments towards a more flexible architecture capable of supporting multiple data sources and destinations. This creates an opportunity to work on a technically challenging environment where Big Data, Cloud, CDC, and modern data engineering practices come together.


Your responsibilities will include

  • Developing and maintaining Change Data Capture (CDC) applications and data ingestion processes in Databricks
  • Creating and maintaining Databricks notebooks for setup, configuration, and data processing
  • Developing and parametrizing batch processing pipelines using Apache Spark
  • Implementing CDC logic to correctly detect and process inserts, updates, and deletes
  • Ensuring the generation of complete data snapshots while maintaining historical changes and data traceability
  • Working with Delta Lake to support reliable, scalable, and performant data processing
  • Working with multiple data sources, including Azure File Share, DB2, SQL Server, and SingleStore
  • Integrating processed data with destinations such as SingleStore and Kafka Confluent
  • Working with Databricks Workflows and Azure services to orchestrate and operationalize data pipelines
  • Optimizing processing performance and ensuring that pipelines are scalable and resilient
  • Contributing to the definition and evolution of data ingestion architecture and processing frameworks
  • Collaborating with other Data Engineers and technical teams to integrate processed data with downstream microservices
  • Ensuring data integrity, consistency, and traceability throughout migration and processing flows
  • Applying best practices around ETL/ELT, data engineering, monitoring, error handling, and pipeline reliability
  • Contributing to continuous improvement initiatives across the data platform and its ingestion frameworks


What should you bring to the team

  • 3+ years of experience in Data Engineering, Databricks, or related data-intensive roles
  • Solid hands-on experience with Databricks, including notebooks, clusters, jobs, and workflows
  • Strong knowledge of Apache Spark, including batch processing and distributed data transformations
  • Strong experience with Python and SQL
  • Hands-on experience with Delta Lake and modern lakehouse architectures
  • Practical understanding of Change Data Capture (CDC) concepts and implementation patterns
  • Experience developing and parametrizing data pipelines and processing frameworks
  • Experience working with cloud data platforms, particularly Microsoft Azure
  • Familiarity with Azure services such as Azure File Share and Data Lake
  • Experience integrating data pipelines with Kafka, ideally Kafka Confluent
  • Good understanding of data architecture and ETL/ELT best practices
  • Experience designing scalable and resilient data processing solutions
  • Strong problem-solving skills and ability to troubleshoot distributed data processing environments
  • Strong communication skills and ability to collaborate with different technical teams
  • Availability to work in a hybrid model, 3 days per week in Tagus Park
  • Fluent Portuguese and English, both written and spoken


Nice to have

  • Experience with SingleStore
  • Experience with SQL-based CDC implementations across different database technologies
  • Experience with file-based CDC or ingestion frameworks
  • Knowledge of microservices and event-driven architectures
  • Experience with MLOps or Data Engineering in cloud environments
  • Familiarity with AI/GenAI pipelines and data platforms supporting AI initiatives
  • Experience with streaming technologies and event-driven data architectures
  • Knowledge of performance optimization techniques for Spark and Databricks
  • Experience designing resilient and highly scalable data solutions


What is in it for you

  • An amazing informal culture of smart, hardworking, and friendly people who support and care about each other
  • A mentorship program from day one
  • Opportunity to work on an innovative project combining CDC, Big Data, Cloud, Databricks, and Kafka
  • The chance to contribute to the evolution of a modern and scalable data architecture
  • Formal training and certifications in data, cloud, and engineering excellence
  • Access to cutting-edge tools and technologies
  • A dynamic company where your ideas matter more than job titles
  • Flexibility with responsibility
  • A Rubber Duck to help you debug those tricky Spark transformations


Don’t hold off any longer and apply now!

If you have any questions, drop us a line at [email protected]

Posted Today

Related Jobs

Related Searches

Apply Now