Search for More Jobs
Get alerts for jobs like this Get jobs like this tweeted to you
Company: AMD
Location: Santa Clara, CA
Career Level: Mid-Senior Level
Industries: Technology, Software, IT, Electronics

Description



ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we're looking for talent who feel the same: people who want to leave the planet better than they found it, those who don't shy away from humanity's challenges but are determined to help solve them.

 

AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you're designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.



Job Role and Responsibility: AMD, Inc., is hiring PMTS Systems Design Engineer to Research, design, develop, and/or test operating systems for semiconductor operations, applying principles and techniques of computer science, engineering, and mathematical analysis. Develop software systems and tools in support of design, infrastructure and technology platforms, including operating systems, compilers, routers, networks, utilities, databases, cloud-based and Internet related tools. Integrate state-of-the-art software solutions that pave the way for ethernet to be used as a viable network technology for the GPU-to- GPU communication required during AI inferencing and training Collaborate with hardware and software teams to enhance the overall performance of GPU clusters, focusing on aspects such as RDMA throughput, latency, and collective communications Provide root cause analysis guidance to internal design teams and lead debug efforts for findings to identify root cause and resolution.  Develop and execute comprehensive benchmarking strategies to assess baseline performance, analyze bottlenecks, and identify areas for improvement within GPU cluster environments. Improve debug capabilities and methodologies by identifying common challenges or impediments to efficient debug and drive innovation in tools and methods. Determine hardware compatibility and/or influence hardware design. Work in an area of specialization to develop systems-level software, working on problems of complex scope where analysis of situations or data requires a review of a variety of factors. Utilize knowledge of computers and electronics, including computer hardware and software, applications, and programming, as well as knowledge of the practical application of engineering science and technology.  Utilize knowledge of computers and electronics, including circuit boards, processors, chips, and electronic equipment, as well as knowledge of design techniques, tools, and principles. Apply knowledge of engineering principles, best practices, and technologies to the design, development, and testing of various AMD semiconductor systems and products.  

 

Can work remotely.

 

Multiple openings.  Qualified applicants click “APPLY NOW” button to apply online.

 

Travel required: NO    

Qualifications: Degree required

Master's degree or foreign equivalent in Computer Science, Computer Engineering, Electrical Engineering, Telecommunications, or related field.

Qualifications: Amount and type of experience required: Five (5) years of experience in the job offered or closely related engineering role.

Alternate combination of education and experience: Employer will alternatively accept a Bachelor's degree or foreign equivalent in Computer Science, Computer Engineering, Electrical Engineering, Telecommunications, or related field and seven (7) years of progressive post-baccalaureate experience in the job offered or closely related engineering role.

 

Specific skills required: The following skills are required:

 

Position requires five (5) years of experience in the following:

 

  • Performance optimization of GPU clusters;
  • GPU architectures, parallel computing concepts, and network protocols;
  • Scripting languages, including Python and Bash, for automation and performance analysis;
  • System level performance analysis tools and methodologies for GPU clusters;
  • Software debugging;
  • Cluster management tools and systems;
  • RDMA network configuration, troubleshooting, and performance tuning;
  • Linux kernel networking; and
  • Machine learning or HPC system design.

 

Can work remotely.

 

Preferred Skills: N/A

 

#LI-AM4



Benefits offered are described:  AMD benefits at a glance.

 

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.

 

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD's “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.


 Apply on company website