Back to Search Results
Get alerts for jobs like this Get jobs like this tweeted to you
Company: AMD
Location: Bayan Lepas, Penang, Malaysia
Career Level: Mid-Senior Level
Industries: Technology, Software, IT, Electronics

Description



ADVANCE YOUR CAREER. ADVANCE THE WORLD. 

At AMD, we believe technology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. 

 

Whether you're designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we'll advance your career.



THE ROLE:  

We are seeking an Infrastructure Platform Engineer with strong MAAS, operating-system imaging, Linux, and bare-metal provisioning experience to provide L2 technical and incident-management support for AMD Fleet Services. This offshore role is an escalation point for incidents that exceed L1 runbooks and is responsible for advanced diagnosis, coordinated restoration, authorized remediation, recovery validation, and complete escalation to the accountable platform or engineering owner.

 

The engineer will improve the reliability and repeatability of server onboarding, reprovisioning, image deployment, host recovery, and Linux lifecycle operations while helping the IOQ Modern NOC convert recurring issues into monitoring, automation, SOP, and runbook improvements.

 

THE PERSON:  

You are a hands-on Linux and infrastructure platform engineer who troubleshoots from evidence, understands the complete bare-metal provisioning path, and remains disciplined during high-impact incidents. You balance restoration speed with change control, communicate ownership and risk clearly, and collaborate effectively across offshore and United States-based teams.

 

KEY RESPONSIBILITIES:  

  • Provide L2 incident management and technical support for in-scope MAAS, Linux, operating-system imaging, bare-metal provisioning, and host-lifecycle services used by AMD Fleet Services.
  • Own technical investigation of assigned incidents from L1 escalation through diagnosis, mitigation, recovery validation, documentation, and handoff or closure.
  • Assess incident impact, priority, affected assets, dependencies, recovery options, risks, and required owners; provide concise updates during incident bridges and follow-the-sun handoffs.
  • Validate L1 evidence, isolate provisioning, image, operating-system, network-boot, hardware-management, or automation failure domains, and execute approved recovery actions.
  • Troubleshoot MAAS controllers, regions, racks, commissioning, deployment, enrollment, machine state, networking, storage layout, metadata, logs, and API interactions within assigned scope.
  • Troubleshoot PXE, DHCP, DNS, TFTP, HTTP, cloud-init, curtin, image synchronization, repository, certificate, proxy, and authentication dependencies that affect provisioning and imaging.
  • Support the creation, validation, publication, and lifecycle management of approved Linux images, ensuring version control, traceability, security requirements, and repeatable deployment outcomes.
  • Diagnose Linux boot, kernel, driver, package, filesystem, service, permission, configuration, performance, and logging issues that prevent host readiness or reliable operation.
  • Use BMC, Redfish, firmware, BIOS, hardware inventory, and out-of-band telemetry to distinguish hardware, firmware, provisioning, and operating-system failures.
  • Coordinate restoration and escalation with network, storage, compute, identity, security, facilities, platform engineering, and hardware support teams while maintaining clear IOQ scope and ownership boundaries.
  • Create complete escalation packages with impact, timeline, affected systems, evidence, actions attempted, change references, residual risk, and recommended next steps.
  • Participate in postmortems and problem-management reviews; identify recurring failure modes and track corrective actions to the accountable owner.
  • Develop and maintain SOPs, runbooks, validation checks, troubleshooting guides, and knowledge articles that enable safe and consistent L1 execution.
  • Automate repeatable diagnostics, evidence collection, image validation, provisioning checks, configuration verification, and approved remediation using Python, shell, Ansible, MAAS APIs, or comparable tools.
  • Improve observability for provisioning and Linux platform services by defining actionable telemetry, alerts, dashboards, service-health indicators, and diagnostic evidence requirements.
  • Maintain accurate Jira records, operational documentation, shift handoffs, and service-status updates throughout the incident lifecycle.

PREFERRED EXPERIENCE:  

  • Experience supporting MAAS or comparable bare-metal provisioning platforms in a production data center, GPU, HPC, private-cloud, or large-scale Linux environment.
  • Strong Linux administration and troubleshooting across boot, kernel, drivers, services, filesystems, networking, permissions, packages, performance, and logs.
  • Hands-on experience with PXE, DHCP, DNS, TFTP, cloud-init, curtin, image repositories, and automated operating-system deployment.
  • Experience building, testing, versioning, securing, and maintaining standardized Linux images and associated deployment pipelines.
  • Experience with server hardware, firmware, BIOS, BMC, Redfish, out-of-band management, and hardware-health telemetry.
  • Automation experience with Python, shell, Ansible, REST APIs, Git, CI/CD, or configuration-management systems.
  • Experience with Jira, major-incident response, change control, postmortems, and problem management.
  • Ability to work independently offshore and provide complete handoffs to United States-based teams.

ACADEMIC CREDENTIALS:

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field preferred; equivalent relevant experience considered.

LOCATION:

Penang, Malaysia

 

#LI-KL1

#LI-Hybrid



Benefits offered are described:  AMD benefits at a glance.

 

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.

 

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD's “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.


 Apply on company website