ML Engineer
Remote (United States)
Job Details
Location: United States
Workplace: Remote
Employment Type: Full-time
Experience: 3+ years developing, training, and deploying deep learning models; 3+ years working with geospatial data processing
Core Areas: Geospatial Foundation Models, Deep Learning, Remote Sensing, MLOps, PyTorch, Geospatial Data Pipelines, Airflow, Model Deployment
Schedule: Pacific or Mountain time zone preferred
Compensation: $100,000 – $200,000 per year
About the Role
This opportunity is for an ML Engineer responsible for building, adapting, and operationalizing foundation model-based deep learning systems that estimate forest structure metrics from remotely sensed data. The role includes fine-tuning geospatial foundation models, developing custom deep neural network heads, preparing training data, evaluating model performance, and integrating trained models into automated production pipelines.
The position combines machine learning, remote sensing, geospatial data engineering, and MLOps. Work spans satellite and lidar datasets, containerized inference services, STAC infrastructure, Airflow DAGs, automated retraining, model monitoring, and scientific documentation while collaborating across science, data engineering, and product teams.
What You'll Do
ML Model Development & Adaptation- Adapt and fine-tune custom or publicly available geospatial foundation models as backbone architectures for domain-specific deep neural network heads that estimate forest structure metrics such as canopy height, biomass, and basal area.
- Prepare, curate, and manage training datasets from Sentinel-2, Sentinel-1, Landsat, lidar, NAIP, and field plot inventories.
- Evaluate model performance using standard remote sensing accuracy metrics and field-based validation data.
- Contribute to experiment design, hyperparameter optimization, and ablation studies in coordination with technical ML leadership.
- Integrate trained machine learning models into automated geospatial data pipelines as containerized and orchestrated inference services.
- Build and maintain STAC (SpatioTemporal Asset Catalog) infrastructure for discovery, cataloging, and access control of ML model inputs and outputs.
- Design and implement larger production pipelines composed of multiple Airflow DAGs while maintaining idempotency, observability, and fault tolerance.
- Maintain and improve data ingestion, preprocessing, and quality-control workflows for satellite imagery and ancillary datasets.
- Monitor pipeline health and model drift, implementing alerting and automated retraining triggers when needed.
- Develop model cards that summarize modeling methods and model performance.
- Write and contribute to scientific manuscripts describing modeling methods, validation results, and novel applications.
- Serve as a cross-functional link between scientific development, data engineering, and product teams by translating requirements, communicating technical constraints, and aligning priorities.
- Document data pipelines, model architectures, and operational procedures in team knowledge bases.
- Participate in code reviews, architectural discussions, and sprint planning.
- Work collaboratively across interdisciplinary teams spanning science, engineering, and product.
- Maintain organized workflows, high-quality data, and clear technical documentation.
- Manage time effectively and work independently in a remote-first environment.
- Communicate effectively across scientific and engineering audiences.
- Support an inclusive and equitable working environment that values diverse perspectives and backgrounds.
- Complete required information-security training, safeguard customer and organizational data, protect credentials, and report suspected security incidents or policy violations through established channels.
- Follow secure development practices and established change-management processes for production systems, protect the confidentiality and integrity of customer data, and promptly address security vulnerabilities within the role's area of responsibility.
Qualifications
Required Experience
- 3+ years of experience developing, training, and deploying deep learning models.
- 3+ years of experience processing geospatial data.
- Experience building and maintaining data pipelines with workflow orchestration tools such as Airflow, Prefect, Dagster, or equivalent platforms.
- Experience with containerization using Docker and familiarity with cloud platforms, with AWS preferred.
Required Skills
- Strong proficiency in Python and the data science stack, including NumPy, pandas, xarray, and scikit-learn.
- Strong deep learning skills, with PyTorch preferred.
- Hands-on proficiency with geospatial processing tools such as rasterio, GDAL, geopandas, and shapely.
- Proficiency with Git, GitHub, collaborative code review, and CI/CD practices.
- Familiarity with STAC specifications and geospatial data catalog infrastructure.
- Strong written communication skills with the ability to contribute to scientific manuscripts and technical documentation.
- Basic knowledge of forest ecology, remote sensing principles, or natural resource science.
Education
- Master's degree in Computer Science, Machine Learning, Remote Sensing, Data Science, Ecology, or a related quantitative field, or equivalent work experience.
Preferred Qualifications
- Ph.D. in a relevant field.
- Experience with geospatial foundation models and self-supervised learning.
- Experience with Kubernetes and distributed computing for large-scale inference.
- Familiarity with ML experiment tracking tools such as MLflow or Weights & Biases and with model registry practices.
- Experience with PostgreSQL, PostGIS, and message queues.
- Publications in remote sensing, machine learning, or ecology journals.
Benefits
- Health, dental, and vision insurance.
- 401(k) plan.
- Unlimited PTO.
- Company equity.
- Cell phone stipend provided each pay period.
- One-time home office setup allowance.
Looking for more opportunities?
View All Jobs