Lead QA Engineer
Remote (United States)
About the Role
Location: United States
Workplace: Remote
Employment Type: Full-Time
Experience: 7+ years in Software QA Engineering, including 2+ years testing AI or ML systems
Core Areas: AI/LLM Evaluation, Test Automation, Python, CI/CD Quality Gates, API and Integration Testing, Regression Testing, Performance and Load Testing, SOC 2 Type II Compliance
Compensation: $140,000 - $165,000 per year plus equity
Travel: Occasional travel to client sites and team offsites
About the Role
This opportunity is for a Lead QA Engineer who will build and establish a formal quality assurance function for a growing AI software environment. This is a hands-on individual contributor role focused on writing test cases, building AI evaluation frameworks, developing automated test coverage, implementing CI/CD quality gates, establishing regression baselines, and triaging production issues alongside Engineering.
The role combines software QA engineering with AI and ML evaluation, release reliability, and regulatory quality requirements. The successful candidate will work with Python, pytest, Playwright or Selenium, REST APIs, Azure, AI evaluation tooling, automated test pipelines, and SOC 2 Type II documentation. Once the QA practice is stable and documented, the position will help bring in Quality Engineering talent to scale the function.
What You'll Do
AI Evaluation- Design and execute evaluation frameworks for LLM and agentic AI outputs across AI development environments and client-deployed instances.
- Write assertions, define behavioral contracts, and establish regression baselines for model behavior.
- Develop evaluation approaches for non-deterministic AI systems where traditional deterministic assertions are insufficient.
- Apply distributions, confidence intervals, and appropriate evaluation methods to complex AI outputs.
- Build tooling that supports reliable evaluation of AI-generated results, including complex outputs such as document summaries.
- Write and maintain automated test suites covering end-to-end, integration, and regression scenarios.
- Build and maintain automated coverage for backend APIs, document ingestion pipelines, AI inference workflows, and frontend surfaces.
- Own performance and load testing for latency-sensitive AI inference paths.
- Implement and enforce quality gates within CI/CD pipelines.
- Participate directly in production bug triage alongside Engineering.
- Create reproducible test cases for production defects and investigate failures across application and infrastructure layers.
- Build regression tests that prevent resolved production issues from recurring.
- Produce test artifacts, audit logs, and QA process documentation that meet SOC 2 Type II standards.
- Maintain defensible testing evidence and quality documentation suitable for scrutiny in a regulated banking environment.
- Work directly with Forward Deployed Engineering on client-side validation.
- Reproduce production issues affecting client environments and validate resulting fixes.
Qualifications
Required Experience
- 7+ years of professional experience in Software QA Engineering.
- At least 2 years of hands-on experience testing AI or machine learning systems.
- Experience writing test cases against LLM outputs and evaluating non-deterministic AI behavior.
- Proven experience building AI evaluation pipelines from scratch.
- Experience building and maintaining automated test suites in production software environments.
- Experience integrating QA quality gates into CI/CD pipelines and owning the process end to end.
Required Skills
- Fluency in Python for test automation, evaluation tooling, and quality engineering workflows.
- Hands-on experience building automated test suites using pytest, Playwright, Selenium, or comparable testing frameworks.
- Hands-on experience with AI evaluation tools such as RAGAS, DeepEval, LangSmith, or comparable tooling.
- Ability to distinguish flaky automated tests from genuinely non-deterministic system behavior.
- Strong understanding of REST APIs and the ability to test and troubleshoot API-driven applications.
- Working knowledge of Azure and asynchronous systems, with the ability to trace failures from the application layer through infrastructure independently.
- Experience with end-to-end, integration, regression, performance, and load testing.
- Ability to establish regression baselines and reliable quality gates for AI and traditional software workflows.
- Ability to produce defensible QA documentation, test artifacts, audit logs, and process documentation that support SOC 2 Type II requirements.
- Ability to work as a hands-on individual contributor while building and documenting a scalable QA practice.
Preferred Qualifications
- Experience working in fintech, banking, or another regulated software environment.
- Familiarity with document processing pipelines.
- Familiarity with multi-agent architectures.
- Experience with Retrieval-Augmented Generation (RAG) validation.
- Familiarity with observability tooling such as Arize or Langfuse.
Success Measures
- Within the first 90 days, complete a diagnostic assessment of current test coverage and share the findings with Engineering leadership.
- Within the first 90 days, build and deploy an evaluation framework for at least one AI-powered workflow.
- Within the first 90 days, implement live quality gates within CI/CD.
- Within the first six months, establish regression baselines for model behavior.
- Within the first six months, document SOC 2 test artifacts and ensure they are audit-ready.
- Within the first six months, establish automated test execution on every release without manual intervention.
- Within the first year, help staff and scale the Quality Engineering function.
- Establish test coverage that scales with product releases and makes quality a first-class input to deployment decisions.
Looking for more opportunities?
View All Jobs