← All roles
TRM Labs logoTRM LabsSecurity

Senior Infrastructure Engineer

NodeGoPythonKubernetesAWSGCPPostgresRemote · Senior · Seed

Build a Safer World.

TRM Labs provides AI-powered intelligence solutions that help public and private sector agencies investigate and disrupt crime. TRM's platforms enable investigators to trace illicit activity, build cases, and construct operating pictures of threat networks. Leading agencies and businesses worldwide rely on TRM to make the world safer and more secure.

Remote-first (US) · Infrastructure Engineering · flexes to Staff for the right builder

In one line

Our team builds the platform that every TRM product, every internal AI agent, and every investigation our customers run depends on. We are not maintaining a system someone else finished - we are building its next generation, on the bleeding edge, for workloads where reliability and security are not negotiable. We're looking for an engineer who has taken infrastructure from zero to one and can operate it in production.

Why this work matters

TRM's platform lets financial institutions and governments trace illicit finance, disrupt crypto- and AI-enabled crime, and run some of the highest-stakes investigations in the world. Our infrastructure is what lets investigators and banks trust the data underneath them. When we get it right, a case moves and a threat gets stopped. When infrastructure is weak, everything above it is slow, fragile, or unsafe. That's the job: make the foundation something the mission can be built on.

What you'll actually work on

This is real, current work - not a hypothetical roadmap:

  • Multi-region, multi-cloud Kubernetes across US GCP, EU GCP, and AWS GovCloud - making "stand up another site" a self-serve configuration instead of a bespoke project.

  • A genuinely self-serve internal PaaS. ~60+ engineers already deploy on it, the large majority with no infra person in the loop. You extend the paved road - golden paths, guardrails, and interfaces engineers actually enjoy using.

  • Agent-native infrastructure. Secure sandboxes where AI agents build and ship, and encoding our operational knowledge into skills so the platform increasingly runs and heals itself.

  • Zero-trust secrets with Vault. Short-lived, dynamic credentials replacing every long-lived secret in the system - closing the blast radius on leaked passwords, one database at a time.

  • Reliability engineering that means it. SLOs and error budgets, real observability, and disaster recovery with validated cross-region restore and sub-hour RTO.

  • Databases as a first-class platform concern. Postgres on Kubernetes, migrations, replication, and redundancy for services that can't blink.

  • Compliance as code. SOC2 and FedRAMP posture built into the platform so security is the default, not a scramble before an audit.

What you'll own

You'll own an Area of Responsibility end to end - a slice of the platform that is genuinely yours. You'll define what "great" looks like for it and build the glide path to get there rather than waiting for one; obsess over the metrics that tell you whether it's actually working; and fix root causes so a class of problems disappears instead of working the ticket queue faster. You own the outcome, not just your layer of it.

Who you are

This is a builder's role, and the bar is high:

  • You've taken infrastructure from zero to one - designed and shipped a platform capability that other engineers came to depend on, not just operated one that already existed.

  • You write real software (Go, Python, or similar). You build tooling, automation, and services - you don't only wire up other people's tools.

  • You're deep in GCP and/or AWS - compute, IAM, VPC, networking, storage, load balancing - and comfortable operating across clouds.

  • You've run Kubernetes in production (GKE/EKS) and are up to date with the state of the art for IaC.

  • You can debug low-level problems without hand-holding - a networking issue, a saturated node, a slow query, a failing deploy - and get to the actual cause.

  • You're fluent across the modern platform toolchain: GitOps (Argo/Flux), containers, CI/CD, observability built on SLOs rather than dashboards no one reads.

  • You have strong instincts for security and databases — zero-trust, secret management, systems hardening; provisioning and migrating production databases. SOC2/FedRAMP exposure is a plus.

If you've built platforms that other engineers rely on every day, we want to talk.

What makes people thrive here

High ownership and low ego. Mission-first. Comfortable where speed and rigor have to coexist - something that takes a quarter elsewhere is expected to move in weeks here, without cutting the corners that matter for security and reliability. You see infra, security, data, and product as one connected system, and you turn one-off fixes into reusable paved roads by instinct.

What should give you pause

The scope is real and the playbook isn't fully written - you'll be writing it. Reliability and compliance expectations are high, and mistakes here show up as real incidents for customers and internal teams. If you prefer an environment where everything is defined before you start, this will feel uncomfortable. If that kind of ownership is exactly what you're after, you'll love it.

How we work

Remote-first (US). Small, high-leverage team with a weekly on-call rotation you'll take part in. We're an AI-native team in practice, not in theory - we use agents in our daily workflow and ship the skills that let the platform operate itself.

Life at TRM

We are building a safer world. That promise shows up in how we work every day.

TRM moves quickly. We are a high velocity, high ownership team that expects clarity, follow-through, and impact. People who thrive here are energized by hard problems, experimentation, and continuous feedback. If something takes months elsewhere, it will ship here in days.

Our work sits at the intersection of AI, national security, and fighting crime. The problems are complex, the stakes are real, and the environment evolves quickly. The pace and intensity of the work reflect the importance of the mission. As a result, the way we operate requires a high level of ownership, adaptability, collaboration, and creative problem-solving.

At TRM, you should expect:

  • Priorities and targets to change quickly as we experiment and iterate

  • Work that often requires operating with a high degree of ambiguity

  • A high level of personal ownership and accountability

  • Close collaboration across teams and functions

  • Frequent, high-touch communication

  • Creative problem solving and out-of-the-box thinking

  • A pace that rewards urgency, adaptability, and outcomes

This environment is energizing for people who enjoy building, solving hard problems, and making progress in situations that are not always fully defined. It also requires comfort navigating ambiguity, adjusting course as new information emerges, and maintaining focus and positivity in a fast-moving and intense environment.

We also recognize that this style of operating is not for everyone. If you are primarily optimizing for predictability or a consistently balanced workload, we encourage you to use the interview process to pressure test whether this environment is truly the right fit. We want teammates who thrive here, not just survive here.

At the same time, many people find this work deeply rewarding. If you are excited by meaningful problems, motivated by ambitious goals, and energized by working alongside mission-driven colleagues, there is a good chance you will find TRM to be an exceptional place to grow and contribute. Learn more: Interviewing at TRM: How We Hire and What Success Looks Like

AI Fluency at TRM

AI fluency is a baseline expectation at TRM.

We believe AI meaningfully changes how top performers operate. We expect every team member to use AI to accelerate and reimagine their craft, not just automate surface tasks.

At TRM, AI fluency means you are among the top 10 percent of operators in your function in how you apply AI to:

  • Accelerate repeatable workflows

  • Structure and solve problems

  • Improve output quality

  • Increase speed and leverage

You will be evaluated on applied AI fluency during the interview process.

Leadership Principles

We hire and grow against three leadership principles. They’re the standards for how we operate, treat each other, and make decisions.

  • Impact-Oriented Trailblazer: We put customers first and move with speed, focus, and adaptability. We treat every plan like an experiment – test, ship, measure, and iterate quickly.

  • Master Craftsperson: We care deeply about our craft. We balance speed with high standards, own outcomes end‑to‑end, and invest in getting better everyday.

  • Inspiring Colleague: We add clarity and energy, not noise. We bring humility, candor, and a one‑team mindset — giving and receiving feedback to make the team stronger.

Join our Mission

At TRM we care deeply about our craft. We are looking for individuals who want their work to matter, who experiment with speed and rigor, and who take pride in building a safer world for billions of people. If you’re excited by TRM’s mission but don’t check every box, we encourage you to apply — we hire for slope, judgment, and the will to learn fast.

TRM is a Series C company with $220M in total funding, backed by Goldman Sachs, Bessemer, Y Combinator, Thoma Bravo, and others. Headquartered in San Francisco, TRM operates as a distributed-first company with hubs in Los Angeles, San Francisco, New York, Washington D.C., London, and Singapore.

We collect the information you provide (such as your resume, work history, and contact details) solely for the purpose of evaluating your candidacy for current and future roles at TRM.

If you are located in the European Economic Area, the United Kingdom, or another jurisdiction with applicable data protection laws, you have the right to access, correct, or request deletion of your personal data at any time before that period ends. To exercise any of these rights, contact us at privacy@trmlabs.com.

To notify TRM Labs that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

The use of AI tools of any kind (including but not limited to notetakers, interview assistants, and real-time coaching tools such as Otter.ai, Fireflies, Fathom, Cluey, or similar) during TRM interviews is not permitted without prior approval from TRM. TRM uses its own internal tools for note-taking to ensure a consistent and confidential experience for all candidates.

Learn More: Company Values | Interviewing | FAQs

AI

Check your CV against this role

Drop your CV. You get a 0-100 fit score against the actual job description, plus the read a senior engineering lead would write. Private to you.

Your CV joins the pool too, so roles that fit can find you. No spam, and nothing reaches a company without your go-ahead.

Score this once, or every future role

Start the candidate journey and every new role on the board gets scored against you.

Five minutes. Tell us what you’re after, drop your CV once, pick how we should reach out. You get a candid read back and you only hear from us when a role fits.

More at TRM Labs