Production Systems Engineer Intern
MetaAdded 3h ago
Save this job and browse the whole board.
No credit card needed
AI analysis
ProJev · full postingQuick summaryby newgrad.ai
What you'll do
- Support Meta's custom servers throughout their 5+ year lifespan in data centers
- Investigate server failures and implement lasting solutions
- Work on hardware root cause analysis, automation, diagnostics, or anomaly detection
What they're looking for
- Interest in production systems engineering and hardware reliability
- Problem-solving skills for complex infrastructure challenges
Pay and perks
- 12-week internship in Menlo Park, CA
We're part of Hardware Design and Release to Production (HDRTP) in Meta's Infrastructure organization. Meta designs its own servers, and our team supports them for the 5+ years they run in Meta's data centers. That includes the compute, database and storage servers that run Facebook, Instagram, WhatsApp and Messenger, along with the GPU and AI accelerator systems that train and serve Meta's Muse models. When these servers fail, we find out why and land durable fixes. Engineers on the team specialize in one of several areas: in-depth hardware root cause analysis, automation that makes debugging and triage faster, hardware diagnostics, or datasets and pipelines that detect anomalies earlier. Your internship project will focus on one of them.
Apply on Meta's site.
Apply nowListing wrong or expired, or you're the employer and want it removed? Let us know.
