Software Engineer, Benchmarking
Remote (United States)
Job Details
Location: United States
Workplace: Remote
Employment Type: Full-Time
Experience: 2+ years of professional software engineering experience
Core Areas: AI benchmarking, evaluation infrastructure, benchmark development, LLM evaluations, Inspect, AI provider integrations, experimentation, research collaboration
Schedule: Flexible hours; overlap with UTC-8 (Pacific Time) and UTC (Greenwich Mean Time) is preferred
Travel: Attendance at three annual team retreats is strongly encouraged
Compensation: $150,000-$325,000 per year
About the Role
This opportunity is for a Software Engineer, Benchmarking focused on building, running, and improving infrastructure used to evaluate frontier AI models. The role combines production software engineering with AI benchmark implementation, evaluation tooling, experimentation, and the development of new benchmarks that help researchers, developers, and policymakers better understand AI capabilities.
You will maintain benchmarking infrastructure, integrate with AI providers, adapt existing benchmarks to internal systems, and help design new evaluations. The position also involves close collaboration with researchers, analysts, and engineers to ensure evaluation data is accurate, useful, and effectively incorporated into research products and publications.
What You'll Do
Implement Benchmarks- Implement AI benchmarks within the evaluation infrastructure, primarily using the Inspect library.
- Expand the range of AI capabilities tracked through the existing benchmark suite.
- Improve benchmark infrastructure so new model releases can be evaluated quickly and reliably.
- Integrate with AI providers and configure existing benchmarks to run effectively on internal infrastructure.
- Contribute to the design and development of new AI benchmarks.
- Pitch and prototype original benchmark ideas, experiments, and other evaluation projects.
- Support internal experiments that improve understanding of frontier AI systems.
- Work closely with researchers, analysts, and engineers to ensure evaluation data and outputs are accurate and insightful.
- Help integrate benchmarking results into research products and publications.
- Support rigorous, trustworthy evaluations of AI capabilities on challenging benchmarks.
Qualifications
Required Experience
- More than 2 years of professional software engineering experience building and maintaining complex systems.
- Experience contributing high-quality, robust, and maintainable production code.
- Comfort working deeply within existing codebases and technical infrastructure.
Required Skills
- Strong software engineering fundamentals and the ability to build reliable evaluation infrastructure.
- Ability to generate original ideas for benchmarks, experiments, and technical projects.
- Ability to learn new AI evaluation concepts, frameworks, and tools quickly.
- Strong interest in rigorous evaluation of AI capabilities and producing trustworthy technical results.
Preferred Qualifications
- Hands-on experience running LLM evaluations.
- Familiarity with AI evaluation frameworks such as Inspect.
- Solid understanding of current AI trends and model evaluation practices.
Benefits
- Flexible remote work environment with flexible hours and schedules for most roles.
- Comprehensive health insurance, with additional locally available benefits where applicable.
- Life insurance and pension benefits where applicable.
- Generous paid time off with no specific annual cap and 30 protected days per year.
- Unlimited personal and sick leave.
- Up to six months of parental leave for permanent staff, combining paid and unpaid leave.
- Flexible expense policy covering equipment, productivity tools, and learning and development opportunities, subject to applicable regulations and manager approval.
- Attendance at three annual team retreats is strongly encouraged.
- Applicants should not include a cover letter, photograph, headshot, or personal information unrelated to the role in the application.
Looking for more opportunities?
View All Jobs