Research Engineer - Benchmarks
Job Title: Research Engineer – Benchmarks
Locations: San Francisco, California, or Singapore
Employment Type: Full-time
Work Arrangement: On-site
Salary: $100,000–$200,000 per year
Visa and Relocation: Support available for exceptional candidates
About HUD
HUD is building infrastructure for creating reinforcement-learning training data and evaluations for frontier AI agents. Its platform and marketplace connect this work with frontier AI laboratories, Fortune 500 companies, and innovative startups.
Backed by Y Combinator and a recent Series A from Standard Capital, HUD is working toward an ambitious goal: developing rigorous, domain-independent quality-control and evaluation capabilities for reinforcement-learning and post-training data.
About the Role
HUD is seeking a Research Engineer to design and build high-quality benchmarks for evaluating frontier AI agents on realistic, domain-specific tasks.
You will own benchmark development from initial task definition through implementation, validation, analysis, and technical reporting. Successful benchmarks must be technically rigorous, practically useful, resistant to gaming, and credible to frontier AI laboratories.
This role is ideal for someone who thinks deeply about what evaluations measure, investigates subtle failure modes, and enjoys converting unfamiliar real-world workflows into reliable benchmark tasks.
What You’ll Do
• Design, implement, and maintain internal benchmarks for frontier AI agents
• Work with subject-matter experts to define realistic domain-specific tasks
• Translate real-world workflows into reproducible evaluation environments
• Build infrastructure for running models and agents reliably against benchmark tasks
• Develop scoring methods, metrics, baselines, and experimental protocols
• Analyze benchmark difficulty, reliability, variance, edge cases, and failure modes
• Detect benchmark contamination, shortcuts, gaming, and misleading performance signals
• Validate whether benchmark results correlate with customer needs and real-world performance
• Improve benchmark quality through experimentation, error analysis, and expert feedback
• Write clear technical documentation, benchmark reports, and research findings
• Communicate effectively with colleagues and external partners across time zones
What We’re Looking For
• Proficiency in Python, Docker, and Linux environments
• Strong understanding of AI benchmarks, model evaluations, or agent environments
• Ability to reason from first principles about task design, scoring, reliability, and failure modes
• Experience designing experiments and analyzing quantitative results
• Curiosity and the ability to understand workflows in unfamiliar technical domains
• Strong written and verbal communication skills
• Ability to execute independently without a prescribed roadmap
• Willingness to work on-site in San Francisco or Singapore
Strong Qualifications
• Publications, technical blogs, open-source projects, or independent research involving benchmarks, model evaluations, or failure analysis
• Experience evaluating LLMs, reinforcement-learning systems, or autonomous agents
• Knowledge of probability, statistics, experimental design, and measurement reliability
• Experience creating tasks, environments, harnesses, or datasets for model evaluation
• Familiarity with benchmark contamination, reward hacking, data leakage, and evaluation validity
• Experience building reliable distributed or containerized evaluation infrastructure
• Strong attention to detail and ability to identify subtle inconsistencies
• Early-stage startup experience or evidence of founder-like independent execution
Students and Recent Graduates
Exceptional graduating students and recent graduates are encouraged to apply. Technical aptitude, intellectual curiosity, and demonstrated ability matter more than a specific number of years of experience.
Strong applicants may demonstrate their qualifications through research, internships, publications, technical writing, GitHub projects, benchmark development, agent environments, evaluation frameworks, or other substantial independent work.
Why Join
• Work on foundational infrastructure for reinforcement learning and frontier AI
• Own technically important benchmarks from concept through deployment
• Collaborate with domain experts and frontier AI organizations
• Investigate difficult questions about model capability, reliability, and failure
• Take meaningful ownership within a small, fast-moving team
• Join a Y Combinator-backed company during a significant stage of growth
• Choose between on-site opportunities in San Francisco or Singapore
Application Requirements
Please submit:
• An updated résumé
• Links to relevant publications, technical blogs, GitHub repositories, or independent projects
• A brief explanation of your interest in AI benchmarks and evaluations
• An example of an evaluation, benchmark, environment, or technical system you have built or studied
• Your preferred location and earliest available start date
For further inquiries, please contact careers@workcentral.ai.