Senior / Staff Infrastructure Engineer (Compute)
About FluidStack
Fluidstack is the AI Cloud Platform. We build GPU supercomputers for top AI labs, governments, and enterprises. Our customers include Mistral, Poolside, Black Forest Labs, Meta, and more.
Our team is small, highly motivated, and focused on providing a world class supercomputing experience. We put our customers first in everything we do, working hard to not just win the sale, but to win repeated business and customer referrals.
We hold ourselves and each other to high standards. We expect you to care deeply about the work you do, the products you build, and the experience our customers have in every interaction with us.
You must work hard, take ownership from inception to delivery, and approach every problem with an open mind and a positive attitude. We value effectiveness, competence, and a growth mindset.
About the Role
We are looking for an Senior / Staff Infrastructure Engineer (Compute) to design, deploy, and manage the compute infrastructure powering Fluidstack's GPU clusters. You will be responsible for ensuring the performance, scalability, and reliability of our compute resources, working closely with hardware and software teams to support our AI workloads.
Focus
Design and implement GPU/ASIC infrastructure at the server, rack, and system level.
Troubleshoot complex GPU and compute system related failures.
Develop and maintain hardware/firmware management services.
Automate all aspects of the server lifecycle.
Own end-to-end compute lifecycle, including partnering with vendors on RMAs.
Serve as the main point of contact for hardware escalation and troubleshooting.
Monitor system performance, identifying and resolving bottlenecks.
Automate deployment and management tasks to improve efficiency.
Collaborate with storage and network teams to ensure cohesive infrastructure operations.
About You
5+ years of experience in compute infrastructure engineering.
Strong knowledge of Linux systems administration and performance tuning.
Experience with bare metal provisioning tools (MaaS, Metal3, Tinkerbell, or other).
Familiarity with GPU hardware and workload optimization, especially kernel and driver level requirements.
Proficiency in automation tools (e.g., Ansible, Terraform).
Experience operating Kubernetes and SLURM clusters.
Benefits
Competitive total compensation package (salary + equity).
Retirement or pension plan, in line with local norms.
Health, dental, and vision insurance.
Generous PTO policy, in line with local norms.
Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
Recommended Jobs
Heavy Equipment Operator (Snow)
We are looking to hire equipment operators for the upcoming 2025/2026 snow season. $30-$45/hr. Seasonal/shift and full-time positions available. Flexible hours and availability. Newest equipm…
Accounts Payable Specialist
About Us: Spear Physical and Occupational Therapy is the nation’s leading outpatient practice. With more than 60 clinics in the New York Tri-State Area and 25 years of experience, Spear provides u…
Senior Backend Software Engineer
As a Senior Backend Software Engineer at Regard, you’ll be involved in all stages of the product development and deployment lifecycle: participate in idea generation, planning, design, prototyping, e…
Lead Product Manager, Risk/LTV
Flex is a growth-stage, NYC headquartered FinTech company that is creating the best rent payment experience. It’s hard to believe that it’s 2025 and paying rent on time is expensive, inflexible, and …
Post-Acute Cardiology Nurse Practitioner in Bronx, NY
Join TeamHealth's growing post-acute care team in the Bronx, New York, area. This is an excellent opportunity to provide quality, compassionate care as a cardiology nurse practitioner (NP) for weekda…
Operations Manager
Job Description Job Description We are a lean size medical imaging service provider located in Queens, NY. The company serves a diverse group of clients to provide quality imaging services. The c…
Logistics Client Support Representative (Korean Bilingual)
OEC Group offers hybrid work, competitive salary, a full benefits package, opportunities for professional growth and so much more! What we’re looking for… Candidates with 1+ years of import frei…
Events Producer
About Us Lox Club’s mission is to bring back the magic of third spaces. We work from home, scroll tiktok for 7 hours a day, and somehow live without the deeply needed mental health benefits our gra…