Member of the Technical Staff - Systems ML Engineer

Transfyr
Transfyr

Software Engineering, IT, Data Science

Cambridge, MA, USA

Posted on Aug 28, 2026

Member of the Technical Staff - Systems ML Engineer

About Transfyr

Transfyr is building physical AI for science.

Why is it that a professional athlete has dramatically more information about every play they make than a scientist has about the cause of any experimental failure? Science has no film room, no instant replay. Instead, a protocol says what was meant to happen. A publication is a lossy record of what might have worked. But all the small decisions and invisible actions that determine whether an experiment succeeds, fails, or transfers to the next lab often disappear the moment the work is done or a scientist leaves.

That missing record is why it has been so hard to automate the physical work of science. It’s why training still remains dependent on scarce, one-to-one apprenticeship. It’s why tech transfer typically requires expensive troubleshooting and is one of the biggest causes of drug launch delays. It’s why scientists struggle to distinguish between biological noise and process variability.

We’re changing that. Transfyr builds physical AI systems that capture real scientific work and turn it into a high-fidelity, machine-readable record of execution and analysis of where process variability is impacting results. In doing so, we are also building the world’s largest commercial dataset on real-world scientific execution. The result is infrastructure that helps teams learn from failures, transfer hard-won know-how, train the next generation of scientists, and give models and robots the grounded data they need to be useful in the real world.

We’re tackling some of the hardest problems at the intersection of frontier science, perception, machine learning, and robotics and have significant traction. We’re backed by a $25M seed round, are collaborating with the largest frontier AI labs, and our advisors include Chris Ré (Stanford), David Baker (Nobel winning UW professor), Kevin Weil (fmr CPO at OpenAI), Steve Quake (Stanford biophysicist), Ken Frazier (fmr CEO of Merck), and Jakob Uszkoreit (CEO of Inceptive and author of “Attention Is All You Need”).

We’re unapologetically ambitious and pragmatic. If you want to work on the hardest problems in the most important industry on earth, join us.

Want to learn more? Read our launch letter here.

The Role

Systems ML Engineers at Transfyr ensure our models train and run efficiently across their full lifecycle, from large-scale training through production inference. You will own performance optimization across our ML stack, working embedded with the research team to make training faster and more efficient, and with our perception and production systems to get models running well at inference, across both cloud and edge infrastructure.

We are looking for engineers who combine deep understanding of ML systems with hands-on performance engineering skill. The ideal Systems ML Engineer has experience profiling and optimizing large-scale models for both training and inference, is comfortable writing custom GPU kernels when off-the-shelf ops aren't fast enough, and can manage cloud and edge infrastructure that holds up in real lab environments.

This role spans deep ML and performance engineering and cloud/DevOps responsibilities. You will profile and optimize training and inference workloads, manage cloud and edge infrastructure, optimize cloud spend, ensure security and compliance, and work closely with the perception and research teams to squeeze more performance out of every model we ship.

We're building a team, and we have needs across levels, from hands-on builders early in their careers to senior engineers who enjoy shaping training infrastructure and technical direction.

This role is in-person in Cambridge, MA.

A tip: While we welcome direct applications, we prefer warm introductions. If you’re really interested in Transfyr, we strongly recommend you get an introduction from someone who knows you well and whose opinion we’re likely to trust. And in general, the strongest way to get our attention is to show us what you have built, solved, or learned that is relevant to this role.

What you'll accomplish with us:

  • Profile and Optimize Performance: Use profiling tools (e.g., Nsight, PyTorch Profiler) to identify bottlenecks in data loading, gradient computation, and communication, and implement optimizations like kernel fusion, sharding, and tiling to improve step time.

  • Optimize Distributed Training: Improve the efficiency of distributed training pipelines using frameworks like PyTorch Distributed, working closely with the research team on training performance.

  • Develop Custom Kernels: Design and maintain high-performance GPU kernels in Triton or CUDA for performance-critical ML workloads.

  • Build Data Pipelines: Design and optimize data loading pipelines that maximize training throughput, and inference pipelines that reliably serve models on real-world, multimodal lab data.

  • Handle Cloud and Edge: Manage deployment across both cloud infrastructure and edge devices running in active lab environments, where compute and connectivity are more limited.

  • Keep Production Running: Debug and resolve performance bottlenecks, resource issues, and failures across the training and deployment stack.

  • Work Closely with Research and Perception: Partner with the research team on training efficiency and with perception engineers to get multimodal data (vision, audio, sensor, metadata) flowing reliably through your pipelines.

  • Make Updates Safe: Build monitoring, versioning, and rollback into deployments so model updates don't break production.

Who you are

  • High agency. You don't wait for perfect datasets or well-posed problems. You identify what needs to be learned, build the right scaffolding, and push work forward.

  • Biased toward action. You prototype quickly, test assumptions against real data, and iterate based on failure rather than waiting for theoretical certainty.

  • Successful in ambiguity. You can make progress when labels are incomplete, feedback is delayed, and success criteria evolve over time.

  • Thoughtful. You understand when sophistication helps and when it obscures, and you make deliberate tradeoffs between model complexity, robustness, and operational cost.

  • Clear, direct communicator. You can explain model performance and limitations to collaborators across engineering, science, and operations.

  • Intense. You care deeply about the mission, work hard when it matters, and help keep the team oriented toward what actually moves the needle.

What you know:

  • Systems ML Engineering Expertise: Demonstrated expertise in ML systems engineering, including optimizing and deploying large-scale models in production, debugging and fixing performance and stability issues in deployed systems, building infrastructure for reproducible, monitored ML deployments, and optimizing inference throughput and resource utilization across cloud and edge.

  • Distributed Training & Serving: Deep knowledge of distributed training and serving frameworks, including PyTorch/JAX distributed strategies, gradient accumulation, mixed precision training, and checkpoint/recovery systems, and how to apply that knowledge to efficient, reliable model serving.

  • Cloud & Edge Infrastructure: Strong cloud administration skills, including AWS services, infrastructure as code (Terraform), Kubernetes orchestration, cost optimization, security best practices, and compliance requirements — plus experience deploying and optimizing ML systems on edge or on-prem infrastructure.

  • Full ML Stack Understanding: Understanding of the ML stack from hardware (GPUs, interconnects, storage) through frameworks (PyTorch, JAX) to deployment and serving, and what it takes to move a model from research prototype to reliable production.

  • Cross-Stack Debugging: Skilled at debugging complex failures across the stack – GPU/NCCL issues, data loading bottlenecks, memory leaks, and performance or convergence problems in both training and deployment.

  • Algorithm Optimization: Deep experience optimizing algorithms for cloud and edge environments, including computer vision and other ML algorithms, with GPU-level work like CUDA and kernel tuning.

Other things we like to see:

  • A passion for and experience in science

  • A passion for and experience with AI

  • Demonstrated experience working in fast-moving/ambiguous environments (like startups!)

The basics:

  • Competitive compensation (cash + equity)

  • Full benefits (low/no-cost health insurance options, HSA, 401K with matching, lunch subsidy, etc.)