Data Engineer (P-175)
SMASH, Who we are?
We believe in long-lasting relationships with our talent. We invest time getting to know them and understanding what they seek as their professional next step.
We aim to find the perfect match. As agents, we pair our talent with our US clients, not only by their technical skills but as a cultural fit. Our core competency is to find the right talent fast.
We purposefully move away from the “contractor” or “outsourcing” type of relationship. Our clients don’t want contractors or “just a service.” Neither does our talent.
This position and offers the opportunity to work with a US-based company. To be eligible for this role, you must have US Citizenship or valid US work authorization.
Role summary
We are looking for an experienced Data Engineer with a strong background in designing, building, and maintaining scalable data pipelines and cloud-based data solutions.
The ideal candidate will bring hands-on expertise with SQL, Python, Spark, Spark SQL, PySpark, and Microsoft Azure, with a strong emphasis on coding and end-to-end pipeline development. Experience with Microsoft Fabric and within the Pharmaceutical, Life Sciences, or Insurance industries will be highly valued.
Responsibilities
- Design, build, and maintain automated data pipelines that move and transform data across systems.
- Develop scalable data workflows to support analytics, reporting, and machine learning use cases.
- Build data ingestion, transformation, processing, and integration solutions within Microsoft Azure.
- Develop and maintain production-quality code using Python and SQL.
- Use Apache Spark, Spark SQL, and PySpark to process and transform large-scale datasets.
- Design data transformations that convert raw data into reliable and usable formats for downstream consumers.
- Develop efficient data integration processes across multiple data sources and destinations.
- Monitor and troubleshoot data pipelines to ensure reliability, accuracy, and performance.
- Identify and resolve data quality, pipeline, and processing issues.
- Optimize data workflows and code for performance, scalability, and maintainability.
- Collaborate with Data Analysts, Data Scientists, engineering teams, and business stakeholders to understand data requirements.
- Support the implementation and continuous improvement of cloud-based data engineering solutions.
- Document pipeline architecture, transformations, dependencies, and technical processes.
- Follow software engineering best practices for coding, testing, version control, and deployment.
Requirements – Must-haves
- 5–6+ years of professional Data Engineering experience.
- Proven hands-on experience designing, building, and maintaining production data pipelines.
- Strong coding and software development capabilities.
- Strong hands-on Python experience.
- Advanced SQL skills.
- Hands-on experience with Apache Spark.
- Strong experience with Spark SQL.
- Strong experience developing data solutions using PySpark.
- Experience building data solutions within Microsoft Azure.
- Experience developing automated workflows for data ingestion, transformation, and delivery.
- Experience processing and transforming large and complex datasets.
- Strong understanding of data integration, ETL/ELT, and data pipeline architecture.
- Ability to troubleshoot and optimize data pipelines and processing workloads.
- Strong understanding of data quality and validation practices.
- Strong analytical and problem-solving skills.
- Ability to work independently while collaborating effectively with cross-functional technical teams.
Nice-to-haves (optional)
- Hands-on Microsoft Fabric experience – strongly preferred.
- Experience building data pipelines or engineering solutions using Microsoft Fabric.
- Pharmaceutical industry experience.
- Life Sciences industry experience.
- Insurance industry experience.
- Experience with enterprise-scale cloud data platforms and distributed data processing.
- Experience supporting data solutions used for analytics, BI, reporting, or machine learning.