Senior Data Engineer at People Data Labs

Summary

Join People Data Labs, a fast-growing global team, and contribute to our mission of democratizing access to high-quality B2B data. As a Data Engineer, you will build and maintain infrastructure for ingesting, transforming, and loading massive datasets using technologies like Spark, SQL, AWS, and Databricks. You will develop CI/CD pipelines and build an entity resolution framework. You'll solve complex data engineering and data science problems and collaborate with stakeholders. This role requires 5-7+ years of experience in data engineering, expertise in Spark and SQL, and strong software development fundamentals. We offer competitive salaries, unlimited paid time off, comprehensive benefits, and the flexibility to work remotely.

Requirements

5-7+ years industry experience with clear examples of strategic technical problem solving and implementation
Strong software development fundamentals
Experience with Python
Expertise with Apache Spark (Java, Scala, and/or Python-based)
Experience with SQL
Experience building scalable data processing systems (e.g., cleaning, transformation) from the ground up
Experience using developer-oriented data pipeline and workflow orchestration (e.g., Airflow (preferred), dbt, dagster or similar)
Knowledge of modern data design and storage patterns (e.g., incremental updating, partitioning and segmentation, rebuilds and backfills)
Experience working in Databricks (including delta live tables, data lakehouse patterns, etc.)
Experience with cloud computing services (AWS (preferred), GCP, Azure or similar)
Experience with data warehousing (e.g., Databricks, Snowflake, Redshift, BigQuery, or similar)
Understanding of modern data storage formats and tools (e.g., parquet, ORC, Avro, Delta Lake)
Must thrive in a fast paced environment and be able to work independently
Can work effectively remotely (able to be proactive about managing blockers, proactive on reaching out and asking questions, and participating in team activities)
Strong written communication skills on Slack/Chat and in documents
You are experienced in writing data design docs (pipeline design, dataflow, schema design)
You can scope and breakdown projects, communicate and collaborate progress and blockers effectively with your manager, team, and stakeholders

Responsibilities

Build infrastructure for ingestion, transformation, and loading an exponentially increasing volume of data from a variety of sources using Spark, SQL, AWS, and Databricks
Building an organic entity resolution framework capable of correctly merging hundreds of billions of individual entities into a number of clean, consumable datasets
Developing CI/CD pipelines and anomaly detection systems capable of continuously improving the quality of data we're pushing into production
Devising solutions to largely-undefined data engineering and data science problems
Work with stakeholders in Engineering and Product to assist with data-related technical issues and support their infrastructure needs

Preferred Qualifications

Degree in a quantitative discipline such as computer science, mathematics, statistics, or engineering
Experience working with entity data (entity resolution / record linkage)
Experience working with data acquisition / data integration
Expertise with Python and the Python data stack (e.g., numpy, pandas)
Experience with streaming platforms (e.g., Kafka)
Experience evaluating data quality and maintaining consistently high data standards across new feature releases (e.g., consistency, accuracy, validity, completeness)

Benefits

Stock
Competitive Salaries
Unlimited paid time off
Medical, dental, & vision insurance
Health, fitness, and office stipends
The permanent ability to work wherever and however you want

Senior Data Engineer

People Data Labs

Summary

Requirements

Responsibilities

Preferred Qualifications

Benefits

Remote

Data

Senior

Similar Remote Jobs

Remote

Data

Senior

Remote

Data

Senior

Netskope

Remote

Data

Senior

Netskope

Remote

Data

Senior

Remote

Data

Senior

Included Health

Remote

Software Development

Senior

United States Department of Defense

Remote

Data

Senior

Wealth

Remote

Data

Senior

LoopMe

Remote

Data

Senior

CoEnterprise

Remote

Sales

Mid-level