Engineer
CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services powering the next wave of AI. Our technology provides enterprises and leading AI labs with the most performant, efficient and resilient solutions for accelerated computing. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe. CoreWeave was ranked as one of the TIME100 most influential companies of 2024.
As the leader in the industry, we thrive in an environment where adaptability and resilience are key. Our culture offers career-defining opportunities for those who excel amid change and challenge. If you’re someone who thrives in a dynamic environment, enjoys solving complex problems, and is eager to make a significant impact, CoreWeave is the place for you. Join us, and be part of a team solving some of the most exciting challenges in the industry. CoreWeave powers the creation and delivery of the intelligence that drives innovation. What You'll Do CoreWeave is seeking a highly skilled and motivated Systems Performance Validation/Tuning Engineer to join our HAVOCK Team, reporting into the Manager of Systems Engineering. In this role, you will play a crucial part in the design, development, and optimization of our bare-metal systems from POST through joining a Kubernetes cluster. The team’s primary responsibilities include maintaining a custom Linux kernel, various OS images (Ubuntu-based), the virtualization stack (kubevirt/qemu/vfio), and the container/pod runtime stack (containerd/nydus/kubelet). You will collaborate closely with cross-functional teams, up stack engineering teams, and stakeholders to ensure our low-level software stack is performant in the context of hardware updates; and providing data, metrics, dashboards, and analysis to substantiate performance assertions. Kernel H ardware - A cceleration - V irtualization - O perating Systems - C ontainerization - K ubelet Our Team’s Stack- Python, Go, bash/sh, C
- Prometheus, Victoria Metrics, Grafana
- Linux Kernel (custom build), Ubuntu
- Intel/AMD/ARM CPUs, Nvidia GPUs, DPUs, Infiniband and Ethernet NICs
- Docker, kubernetes (k8s), KubeVirt, containerd, kubelet
- Develop and maintain tools for establishing systems performance baselines
- Develop and maintain performance regression analysis testing automation
- Design and maintain performance regression test pipelines for HPC workloads
- Debug and Tune fabric-level performance to ensure low-latency high throughput configurations
- Development of telemetry for performance analysis across distributed clusters of servers
- Triage and fix performance issues in Linux
- Collect data, produce metrics and visualizations that communicate performance information compared to benchmarks; this data should lead to appropriate business decisions and toward greater automation that improves customer experience in relation to performance
- Define Linux and OS requirements, specifications, and system architecture in relation to systems performance, in collaboration with cross-functional teams. Along with these responsibilities there will also be cross team collaboration to triage and resolve bottlenecks
- 5+ years of professional experience in Systems/HPC Performance Engineering, Benchmarking, and/or Validation.
- Strong experience with MPI workloads and distributed system performance analysis
- Familiarity with RoCE, InfiniBand, and GPUDirect/Data Direct I/O, NUMA, etc in HPC workloads
- Hands-on use of public HPC benchmarks (HPCC, HPL, OSU, MLPerf-HPC, STREAM, IO500)
- Extensive, deep experience in Linux internals
- Fluency with a programming language geared toward automation (Python preferred, but others possible)
- Experience writing robust, testable code
- Experience diagnosing and fixing systems performance issues
- Experiencing with implementing automation testing
- Ability to effectively prioritize and communicate proposed features and fixes in a remote-employee environment
- Strong passion for automation, with a commitment to automating processes comprehensively
- Excellent documentation skills and attention to detail
- Strong analytical and problem-solving abilities
- Familiarity with QA/QE best practices
- Familiarity with Golang
- Opinions about software version control and team collaboration
- Experience working in Cloud environments
- Experience as a software engineer writing large-scale applications
- Experience in open-source community software development
- Experience with machine learning is a huge bonus
- Be Curious at Your Core
- Act Like an Owner
- Empower Employees
- Deliver Best-in-Class Client Experiences
- Achieve More Together
- Medical, dental, and vision insurance - 100% paid for by CoreWeave
- Company-paid Life Insurance
- Voluntary supplemental life insurance
- Short and long-term disability insurance
- Flexible Spending Account
- Health Savings Account
- Tuition Reimbursement
- Ability to Participate in Employee Stock Purchase Program (ESPP)
- Mental Wellness Benefits through Spring Health
- Family-Forming support provided by Carrot
- Paid Parental Leave
- Flexible, full-service childcare support with Kinside
- 401(k) with a generous employer match
- Flexible PTO
- Catered lunch each day in our office and data center locations
- A casual work environment
- A work culture focused on innovative disruption
- 1157, or (iv) asylee under 8 U.S.C.
- 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.
Recommended Jobs
Email/Chat Support Specialist
This is a remote position. We are seeking a dedicated Email/Chat Support Specialist to join our team at Stello Foods. In this role, you will be responsible for providing exceptional customer servi…
Medical Scribe - Adult Primary Care
Medical Scribe-Adult Primary Care Location: 620 Foster Avenue Brooklyn, NY 11230 Hours: Full Time Sunday, Monday, Tuesday, Thursday: 10:00 AM -6:00 PM Wednesday: 11:00 AM-7:00 PM Pr…
Teaching assistant
Coalition for Hispanic Family Services Job Description Job Title: Teaching Assistant Department: Arts & Literacy - Youth Development Reports to: Program Site Director Date Availab…
Backend Engineer | Financial Systems
About Ramp At Ramp, we’re rethinking how modern finance teams function in the age of AI. We believe AI isn’t just the next big wave. It’s the new foundation for how business gets done. We’re investi…
Med-Surg/Tele Travel RN
Job Summary We’re seeking an experienced and dedicated Registered Nurse (RN) with Medical-Surgical and Telemetry (MedSurg/Tele) experience to join the day shift team in Auburn, New York . Th…
Lead Product Manager, Trading Systems and Infrastructure
About Uphold A fast-growing fintech company, Uphold is pursuing a mission to accelerate access to blockchain technologies for people and companies worldwide. Founded in 2014, Uphold has millions o…
Senior Software Engineer, Data
About Us Camber builds software to improve the quality and accessibility of healthcare. We streamline and replace manual work so clinicians can focus on what they do best: providing great care. Fo…
Engineering Manager
Company Description NBCUniversal is one of the world's leading media and entertainment companies. We create world-class content, which we distribute across our portfolio of film, television, and…
Optometrist
A well-established eye care company is seeking an Optometrist to deliver exceptional vision care to patients in nursing homes throughout the Brooklyn/Queens area of NYC. This position offers flexibil…
Hardware manager
Job Posting information Build Your Career with Curtis Lumber! Founded in 1890, Curtis Lumber is a family owned and operated building materials retailer, one of 100 largest and fastest-growi…