Senior Data Engineer
Remote (United States)
Job Details
Location: United States
Workplace: Remote
Employment Type: Full Time
Experience: 5+ years of data engineering experience
Core Areas: Apache Airflow, dbt, AWS Data Services, SQL, Python, Data Lake & Warehouse Architecture, Data Quality, AI/ML Data Pipelines
Compensation: $190,000-$240,000 per year
About the Role
This opportunity is for a Senior Data Engineer to build, optimize, and operate large-scale data infrastructure supporting multi-state broadband programs. The role is highly hands-on, with significant responsibility for production Apache Airflow DAGs, dbt pipelines, data modeling, AWS-based infrastructure, data lake and warehouse architecture, and reliable ingestion, transformation, storage, and export workflows.
The position also works with time-series and event-based data, automated data quality systems, vector databases, and production AI/ML workflows. The engineer will contribute to architecture and tooling decisions, improve observability and reliability, mentor other data engineers, and help establish engineering standards across the data platform.
What You'll Do
- Partner with product and engineering teams to translate requirements into technical architecture that balances cost, performance, scalability, and long-term maintainability.
- Evaluate and recommend data tools, frameworks, and infrastructure based on technical and product requirements.
- Contribute to roadmap planning and help evaluate build-versus-buy decisions.
- Design, implement, and maintain scalable data infrastructure using AWS services including S3, Athena, RDS/PostgreSQL, ECS, Lambda, Secrets Manager, and CloudWatch.
- Own data lake and warehouse architecture, including partitioning strategies, storage optimization, and data lifecycle management.
- Build and maintain production-grade Apache Airflow DAGs for data ingestion, transformation, and export workflows.
- Implement monitoring, alerting, and observability across the data platform and support rapid incident resolution.
- Build and maintain robust dbt pipelines with data quality checks and well-structured data models.
- Design and maintain database schemas that support multi-state, multi-tenant program data.
- Write and optimize SQL queries across PostgreSQL, Redshift, and Athena for analytical and operational workloads.
- Develop reusable data models, utilities, and shared Python packages for use across the data engineering team.
- Design data models and infrastructure for large-scale time-series and event-based data.
- Implement efficient ingestion, storage, and retrieval patterns for high-frequency temporal datasets.
- Work with vector databases to support AI-powered features and semantic search.
- Integrate LLM workflows into ELT pipelines using AWS Bedrock, LangChain, and related frameworks.
- Build production-grade AI-assisted pipelines for data comparison, validation, and enrichment.
- Work with MLflow and production-level machine learning pipelines.
- Evaluate evolving AI/ML tooling and introduce relevant technologies where they support data platform requirements.
- Design and build automated data QA systems that validate data quality, completeness, and consistency.
- Implement data cleansing and reconciliation processes for complex, multi-source ingestion workflows.
- Proactively monitor data pipelines and resolve issues before they affect downstream consumers or clients.
- Mentor junior and mid-level data engineers through code reviews, pair programming, and architectural guidance.
- Help establish engineering standards for code quality, testing, and technical documentation.
- Promote ownership, technical curiosity, and continuous improvement across the data engineering team.
Qualifications
Required Experience
- 5+ years of data engineering experience, including ownership and operational support of production systems end to end.
- Experience evaluating technical tradeoffs and making pragmatic architecture decisions that balance cost and performance.
- Hands-on experience with Apache Airflow or comparable orchestration tools in production environments.
- Experience working with large-scale time-series, event-based, or streaming data systems.
- Experience using version control systems such as GitHub, Subversion, GitLab, or Mercurial.
- Experience building data quality frameworks or automated validation systems.
- Demonstrated experience mentoring engineers and contributing to engineering team practices.
Required Skills
- Strong proficiency in SQL, including PostgreSQL and Athena/Presto, with the ability to write production-quality queries.
- Strong proficiency in Python and the ability to develop production-quality data engineering code.
- Deep hands-on expertise with AWS data and infrastructure services, including S3, Athena, RDS, ECS, Lambda, IAM, and CloudWatch.
- Strong ability to design and operate production data pipelines, data models, schemas, ingestion workflows, and scalable data infrastructure.
- Familiarity with vector databases and LLM integration patterns using technologies such as LangChain, AWS Bedrock, or similar frameworks.
- Ability to communicate complex infrastructure and architecture decisions clearly to non-engineering stakeholders.
- Ability to work effectively in ambiguous environments and make decisions using first-principles reasoning.
- Ability to collaborate effectively across multiple time zones.
Preferred Qualifications
- Experience managing data for SaaS platforms.
- Experience with large-scale spatiotemporal data pipelines, including PostGIS, spatial indexing, and tiling.
Benefits
- Competitive benefits for employees and dependents.
- Professional learning and growth opportunities.
- Opportunity to contribute to product and technical roadmap decisions.
- Work-from-anywhere flexibility with reliable internet access.
- Optional, encouraged team retreats held 2–3 times per year.
- Opportunity to provide input as the benefits program continues to evolve.
Looking for more opportunities?
View All Jobs