Senior Software Engineer, Data Acquisition

People Data Labs
Summary
Join People Data Labs (PDL), a leading provider of people and company data, as a Data Engineer. You will play a crucial role in building standalone data products, improving existing datasets, and pursuing new ones. This involves using and developing web crawling technologies, supporting infrastructure, structuring data, and developing new techniques to increase efficiency and scalability. You will also build data pipelines, publish data, and work with the data product and engineering team to design and implement new data products. The ideal candidate has 7+ years of industry experience, strong software development skills, experience building crawlers, and proficiency in Linux/Unix. This role offers a high level of autonomy and opportunity for direct contributions within a fast-paced, collaborative environment.
Requirements
- 7+ years industry experience with clear examples of strategic technical problem solving and implementation
- Strong software development architecture and fundamentals for backend applications
- Solid understanding of browser rendering pipeline, web application architecture (auth, cookies, http request/response)
- Solid programming experience: strong grasp of object-oriented design and experience building applications using asynchronous programming paradigms (e.g., async/await, event loops, or concurrency libraries)
- Experience building crawlers
- Proficient in Linux / Unix command line utilities, Linux system administration, architecture, and resource management
- Experience evaluating data quality and maintaining consistently high data standards across new feature releases (e.g., consistency, accuracy, validity, completeness)
- Must thrive in a fast paced environment and be able to work independently
- Can work effectively remotely (able to be proactive about managing blockers, proactive on reaching out and asking questions, and participating in team activities)
- Strong written communication skills on Slack/Chat and in documents
- You are experienced in writing data design docs (pipeline design, dataflow, schema design)
- You can scope and breakdown projects, communicate and collaborate progress and blockers effectively with your manager, team, and stakeholders
Responsibilities
- Use and develop web crawling technologies to capture and catalog data on the internet
- Support and improve our web crawling infrastructure
- Structure, define, and model captured data, providing semantic data definition and automate data quality monitoring for data that we crawl
- Develop new techniques to increase speed, efficiency, scalability, and reliability of web crawls
- Use big data processing platform to build data pipelines, publish data, and ensure the reliable availability of data that we crawl
- Work with our data product and engineering team to design and implement new data products with captured data, and enhance and improve upon existing products
Preferred Qualifications
- Degree in a quantitative discipline such as computer science, mathematics, statistics, or engineering
- Experience as a Red Teamer
- Experience working in data acquisition
- Experience in network architecture and how to debug and inspect network traffic (DNS, IPv4, Proxies, Application ports and interfaces; packet capture and analysis)
- Experience with Apache Spark
- Experience with SQL, including writing advanced queries (e.g., window functions, CTEs)
- Experience with streaming data platforms (e.g. Kafka or other pub/sub; Spark streaming or other stream processing)
- Experience with cloud computing services (AWS (preferred), GCP, Azure or similar)
- Experience working in Databricks (including delta live tables, data lakehouse patterns, etc.)
- Knowledge of modern data design and storage patterns (e.g., incremental updating, partitioning and segmentation, rebuilds and backfills)
- Experience with data warehousing (e.g., Databricks, Snowflake, Redshift, BigQuery, or similar)
- Understanding of modern data storage formats and tools (e.g., parquet, ORC, Avro, Delta Lake)
Benefits
- Stock
- Competitive Salaries
- Unlimited paid time off
- Medical, dental, & vision insurance
- Health, fitness, and office stipends
- The permanent ability to work wherever and however you want