Research Engineer
This is one of 8 engineering roles AfterQuery has open, and one of 6 at mid level. It went up 2 days ago.
About AfterQuery
AfterQuery is an applied research lab curating data solutions for foundation model development. We serve every frontier AI lab with the mission of delivering the best data to power the best models. In doing so, we can make expertise that once took a lifetime to build available to anyone who needs it. Our customers are the ones building the foundation models themselves and our work sits directly in the loop of how those systems improve. This is a rare opportunity to join a company at a defining moment in AI. Since raising our $30M Series A at a $300M valuation, AfterQuery has grown well over a $100M revenue run rate. We're based in San Francisco and backed by leading investors including Altos Ventures, BoxGroup, and Y Combinator and angels from Google DeepMind, OpenAI, Anthropic, Meta Superintelligence Labs, and Microsoft AI.
Why Apply
Massive Opportunity: We are one of the fastest-growing YC companies in our batch, and we believe we can become one of the fastest-growing YC companies of all time.
Founding Impact: You will own and architect core infrastructure systems that power our platform from the ground up.
Equity & Growth: Competitive salary and meaningful equity. As we scale, you’ll have the opportunity to shape the engineering organization and lead major technical initiatives.
Strong Team: Our founding team has experience from Citadel Securities, Meta, Google, Silver Lake, and Morgan Stanley — work alongside world-class engineers and researchers.
Overview
As the Research Engineer, you will design and run training experiments that isolate the impact of our datasets on model behavior. Through controlled SFT and RL - base post-training experiments, you will measure how different data sources, structures, and selection strategies affect capability, generalization, and alignment. You will own the full experiment loop: formulate the hypothesis, build the pipeline, run the model, analyze the results, identify confounders, and determine what should happen next. Working with partner labs and internal teams, you will turn our datasets into clear, defensible evidence: this data -> this improvement -> under these conditions. This is empirical, high-leverage work for someone who likes building quickly and extracting signal from noisy results.
Responsibilities
Design and build our post-training infrastructure for SFT, RL and evaluation workflows.
Build reliable pipelines for data preparation, dataset versioning, sampling, training, checkpoint management, and evaluation.
Develop experiment orchestration and tracking systems that make runs reproducible, comparable, and easy to debug.
Create reusable abstractions that allow researchers to launch experiments quickly across datasets, models, and training recipes.
Integrate our systems with partner-lab training stacks, model APIs, compute environments, and evaluation infrastructure.
Required Qualifications
At least 2 years of professional experience in machine learning engineering, research engineering, ML infrastructure, or a closely related field; 2-4+ years preferred.
Strong Python and software-engineering skills.
Hands on experience with PyTorch, JAX, Ray, and Slurm.
Experience building production-quality ML training, evaluation, or data infrastructure.
HAnds-on experience with LLM fine-tuning, post-training, and evaluation.
Ability to build reliable, reproducible systems for launching and comparing ML experiments.
Strong debugging skills across distributed systems, data pipelines, training infrastructure, and model behavior.
Understanding of experimental design and the ability to extract actionable conclusions from noisy results.
Ability to move quickly between infrastructure engineering and hands-on experimentation.
A bias toward building, testing, and shipping.
Company Benefits (For Eligible Employees):
Health Insurance: Medical, Vision, Dental
401(k) with Employer Match
Daily Meals: Daily UberEats Stipend
Monthly Wellness Stipend
Commute Covered
Check your resume against this role
Drop your resume. You get a 0-100 fit score against the actual job description, plus the reasoning and the gaps worth closing first. Private to you.
Score this once, or every future role
Start the candidate journey and every new role on the board gets scored against you.
Five minutes. Tell us what you’re after, drop your resume once, pick how we should reach out. You get a candid read back and you only hear from us when a role fits.