Freelance Academic Research Expert AI Evaluator
Uber AI Solutions is Uber's marketplace connecting skilled freelancers with Generative AI researchers and product teams working on cutting-edge AI systems. We partner with independent contractors to support large-scale model evaluation, alignment, and quality initiatives.
As a Freelance Academic Research Expert, you will pair genuine scholarly expertise with hands-on GenAI-evaluation experience to design challenging, realistic prompts backed by source documents such as peer-reviewed papers, published datasets, methodology standards, and systematic review protocols. You will then probe where the model breaks, and encode expert judgment into precise, gradable rubrics and a canonical 'golden' answer. You own an end-to-end evaluation task from prompt design through reviewer sign-off.
This is hands-on academic work with an analytical core. You will produce scholarly deliverables as the output; what makes the work difficult is evaluating methodology, evidence, and citation integrity against the provided sources. It is not a teaching, curriculum delivery, or tutoring role.
What you'll work on
- Apply your academic research expertise to help create and evaluate complex AI prompts and responses.
- Conceptualize and draft realistic, multi-step prompts that mirror real research assignments and require synthesis, methodological evaluation, and reconciliation across multiple sources.
- Identify and document genuine model failures by writing specific, objective, and actionable explanations.
- Build comprehensive evaluation rubrics that trace each required input through dependent criteria and carefully separate extraction from interpretation.
- Produce fully correct, deliverable-quality golden responses that satisfy all criteria using discipline-standard conventions and citation practice.
- Complete our onboarding process and strictly follow the disciplined authoring and review workflow for each task.
- Closely follow guidelines, run AI quality checks on your own work, and resolve any issues in the feedback loop before final submission.
Engagement details
- Location: Remote (United States)
- Engagement Type: 1099 Freelance / Independent Contractor, through the Uber AI Solutions platform
- Engagement Duration: 40 hours over 1-2 weeks
Who we're looking for
- Ph.D., Ed.D., or completed Postdoctoral fellowship from an accredited institution.
- 5+ years of experience in academic research, peer review, or higher education instruction.
- Demonstrated ability to read primary source documents (peer-reviewed literature, published datasets, methodology standards) and derive multi-step analytical conclusions from them.
- Familiar with research methodology evaluation, systematic literature review, and citation integrity standards.
- Comfortable working independently on detail-oriented tasks.
- Strong analytical and written communication skills.
- Prior hands-on experience with GenAI / LLM evaluation and prompt engineering, rubric or golden-answer authoring, data annotation, RLHF, red-teaming, or model quality assessment.
Why this matters
Your expertise will help improve how AI systems handle complex scholarly topics, making outputs more accurate, reliable, and better aligned with real-world research reasoning.