← All roles
Wafer logoWaferEngineering, Product and Design
Posted today

Member of Technical Staff

CUDASF · Staff · Seed

Member of Technical Staff

Our mission at Wafer is to maximize intelligence per watt by using AI to optimize AI infrastructure, achieving orders of magnitude better energy and cost efficiency per token.

We believe cheap intelligence is the most essential piece of technology for a future of abundance. We care about building a future where intelligence is "too cheap to meter."

Wafer commercializes these efforts by serving serverless and dedicated inference for open source LLMs at the best performance per dollar. Our core bet is doing this through autonomous optimization of heterogeneous hardware.

What you'll do

  • Ship day-zero support for new open-source models, tuned for latency and throughput

  • Optimize the serving stack: batching, KV cache, speculative decoding, quantization

  • Write and tune kernels in CUDA, HIP, and Triton for NVIDIA, AMD, TPU, Trainium, D-Matrix, and more.

  • Design, deploy, and operate heterogeneous clusters across vendors

  • Run production inference across a mixed fleet: reliability, observability, and cost per token at scale

How we evaluate

We score every candidate on seven values:

  1. Infinitely Resourceful

  2. Exceptionalism

  3. Unreasonable Standards

  4. Company Over Self

  5. High EQ

  6. Learns Quickly

  7. First Principles Thinker

Compensation and benefits

  • $200K base salary + 1–2% equity.

  • Fully covered medical, dental, and vision insurance.

  • Daily lunch and dinner, unlimited PTO, and parental leave.

  • $1K/month housing stipend (post-tax) if you live within walking distance (0.5 miles) from the office.

  • Covered Uber/Waymo from/to office.

  • Visa sponsorship available.

How we work

On-site in San Francisco, five days a week. Small team with massive surface area and ownership. You operate with complete autonomy of how to solve problems, and work with the team to set the direction of your work. We don't see engineers as code writers, but as problem solvers. You will do everything from talking to customers to writing custom GPU kernels in esoteric hardware.

AI

Check your CV against this role

Drop your CV. You get a 0-100 fit score against the actual job description, plus the read a senior engineering lead would write. Private to you.

Your CV joins the pool too, so roles that fit can find you. No spam, and nothing reaches a company without your go-ahead.

Score this once, or every future role

Start the candidate journey and every new role on the board gets scored against you.

Five minutes. Tell us what you’re after, drop your CV once, pick how we should reach out. You get a candid read back and you only hear from us when a role fits.