AfterQuery
ML Research Expert — Published First Author (Benchmark Authoring)
$150 / hourRemote · Worldwide
- Machine Learning
- Benchmark Authoring
- AI Evaluation
Paid, remote work authoring and validating research-grade ML benchmark tasks. For Master's or PhD-level ML researchers with at least one first-author paper.
Domain: Machine Learning / AI Research Requirements
- Experience: 1+ years of hands-on ML research; no upper limit
- Education: Master's or PhD (in progress is fine, including current PhD candidates)
The work
You will design and validate tasks that measure what frontier AI models can actually do across machine learning. We are looking for people who write ML code and publish results, not for annotation or data-labelling experience.
Task areas span: Language Models, Deep Learning, Reinforcement Learning, Vision & Generation, Robotics, ML Systems & Efficient ML, Optimization & Theory, Classical & Adaptive Learning, Time Series & Forecasting, Structured & Causal Reasoning, Trustworthy Learning, and AI for Science.
Application process
Candidates who meet the requirements and provide the requested materials will be prioritised for review. As part of this process we conduct thorough background checks. Please apply only if you meet these requirements.
What you'll do
- Design realistic ML research tasks and problem sets within your area of expertise
- Author expert-level reference solutions and grading rubrics
Requirements
- At least one first-author research paper in machine learning or a closely related field
- Master's or PhD in ML, CS, statistics, mathematics or a related quantitative field, completed or in progress
- Hands-on experience writing ML code (training, evaluation or systems work), not annotation or data labelling
- Meets the experience and education bar described in Full Description
Nice to have
- Multiple first-author publications, or papers reporting measured improvements over a baseline
- Publications at venues such as NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, CoRL or MLSys
- Depth in one or more of the listed task areas rather than broad familiarity
Description from AfterQuery. Full details and the application are on AfterQuery's site.