โšก New

Pyspark Developer

Tata Consultancy Services

ChennaiFull-timeMid LevelOn-site

Job Description

Greetings from Tata Consultancy Services (TCS)!


We are looking for experienced professionals with strong expertise in PySpark, Big Data Testing, Hadoop Ecosystem, and Data Validation to join our data engineering team.


Experience: 5 to 10 Years


Location: Chennai, Kolkata, Hyderabad, and Pune


Technical Skills Required

  • Strong experience in Big Data technologies including Hadoop, PySpark, Scala, Hive, Impala, SQL, and Python.
  • Hands-on expertise in Python scripting and PySpark application development/testing.
  • Experience in Spark UI analysis, debugging, optimization, and performance tuning.
  • Good understanding of SQL concepts including Joins, Subqueries, CTEs, and complex query development.
  • Exposure to database technologies and large-scale data processing environments.
  • Experience working with distributed systems such as HDFS, Hive, and Hadoop clusters.
  • Understanding of ETL/Data Pipeline testing methodologies and data quality validation.


Roles & Responsibilities

  • Design and develop scalable PySpark-based testing frameworks for ETL and data pipeline validation.
  • Architect end-to-end data validation solutions within Hadoop ecosystems, covering schema validation, lineage checks, and data quality monitoring.
  • Design and manage Hadoop/Hive test environments, including YARN resource management and dynamic partitioning.
  • Implement automated testing solutions integrated with Zephyr, Jira, ServiceNow, and API-driven test management tools.
  • Build and maintain CI/CD test pipelines for PySpark and Hadoop applications with artifact management and parallel execution capabilities.
  • Develop data quality validation frameworks using PySpark integrated with Hive metadata services.
  • Create test data generation utilities and reusable testing platforms for large-scale data validation.
  • Configure and optimize Spark sessions, including memory management and executor/core allocation strategies.
  • Work extensively with HDFS and Hive tables for distributed data storage and processing.
  • Implement best practices such as partitioning, caching, and modular transformation pipeline development.
  • Perform Spark performance optimization using techniques such as salting, minimizing shuffles, and execution plan analysis.
  • Mentor junior team members on Big Data testing concepts, automation frameworks, and PySpark testing methodologies.
  • Participate in test strategy discussions and contribute to continuous improvement initiatives.


Preferred Skills

  • Experience with CI/CD tools and DevOps practices.
  • Knowledge of cloud-based data platforms.
  • Strong analytical and troubleshooting skills.
  • Experience in Agile/Scrum development environments.
  • Excellent communication and stakeholder management skills.


Interested candidates can share their updated CV.

Posted Today

Related Jobs

Related Searches

Apply Now