LogoKode$word
Xai logo
Verified Tech Organization

Careers at Xai

Browse and filter through all verified positions currently open at Xai.

Total Company Roles79
Matching Filter79
x.aiHQ: New York, New York, United States

Founded in 2014, x.ai makes an artificial intelligence personal assistant who schedules meetings for you. There's no sign-in, no password, no download; all you do is CC amy@x.ai into your email thread, just like you would a human personal assistant. Amy then takes over the tedious email ping pong that comes along with scheduling a meeting. We're a hardcore technology company, developing invisible software. We build our business sustainably through passionate and loyal customers-and every single team member, scientist or not, has a mission of delivering exceptional customer service at all times. Backed by blue chip investors, including IA Ventures, Firstmark, Two Sigma Ventures, SoftBank Capital, DCM and Pritzker Group, the team is located in New York City.

Sector:aiartificial intelligenceb2bchatbots softwareenterprise software

All Openings (79)

Ordered by most recently published

Spring 2027 Software Engineering Internship/Co-op

On-siteinternshipInternshipPalo Alto, United States
Apply Now

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: SpaceXAI seeks extraordinary students to join us for Spring 2027 software engineering roles. As an intern, you will work closely with your mentor and other employees who will help you apply your knowledge and grow your skills on projects that have a significant impact. If you’ve demonstrated a commitment to academic success and motivation to apply your knowledge outside of the classroom, you are a great candidate! SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. This includes Grok, the frontier AI model for everything you need, trained on the world's largest supercluster, and developing Starmind, a new constellation of satellites capturing solar energy in space to power low-cost, high-performance AI compute for Earth. BASIC QUALIFICATIONS: Must be enrolled in a bachelor’s degree or graduate program 3+ months of software programming or development experience Software coding experience in one or more of the following: C, C++, C#, Java, JavaScript, Python PREFERRED SKILLS AND EXPERIENCE: GPA of 3.5 or above 6+ months experience developing and deploying software that has been used on real-world applications and projects Strong fundamental knowledge of computer architecture and networks Experience with software documentation, creating system diagrams, and enumerating software requirements Strong skills in debugging, performance optimization and unit testing Strong interpersonal skills (examples: leading a student organization or working successfully in teams) Ability to work effectively in a dynamic environment with changing needs and requirements Ability to work independently and in a team, take initiative, and communicate effectively ADDITIONAL REQUIREMENTS: Able to work full time, onsite for a minimum of 12 consecutive weeks beginning in January or March 2027 COMPENSATION AND BENEFITS: Software Engineering Intern/Freshman/Sophomore: $30 USD per hour Software Engineering Intern/Junior/Senior: $34 USD per hour Software Engineering Intern/Completed Bachelor's: $36 USD per hour Software Engineering Intern/Completed Master's: $38 USD per hour Software Engineering Intern/Completed PhD: $40 USD per hour Your salary will be determined by academic level. Hourly pay is just one part of our total rewards package at SpaceXAI. You may also be eligible for a stipend to subsidize relocation costs, as well as access to our comprehensive medical coverage and a 401(k) retirement plan. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .

View more...
Software Engineering
VerifiedToday

Summer 2027 Software Engineering Internship/Co-op

On-siteinternshipInternshipPalo Alto, United States
Apply Now

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: SpaceXAI seeks extraordinary students to join us for Summer 2027 software engineering roles. As an intern, you will work closely with your mentor and other employees who will help you apply your knowledge and grow your skills on projects that have a significant impact. If you’ve demonstrated a commitment to academic success and motivation to apply your knowledge outside of the classroom, you are a great candidate! SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. This includes Grok, the frontier AI model for everything you need, trained on the world's largest supercluster, and developing Starmind, a new constellation of satellites capturing solar energy in space to power low-cost, high-performance AI compute for Earth. BASIC QUALIFICATIONS: Must be enrolled in a bachelor’s degree or graduate program 3+ months of software programming or development experience Software coding experience in one or more of the following: C, C++, C#, Java, JavaScript, Python PREFERRED SKILLS AND EXPERIENCE: GPA of 3.5 or above 6+ months experience developing and deploying software that has been used on real-world applications and projects Strong fundamental knowledge of computer architecture and networks Experience with software documentation, creating system diagrams, and enumerating software requirements Strong skills in debugging, performance optimization and unit testing Strong interpersonal skills (examples: leading a student organization or working successfully in teams) Ability to work effectively in a dynamic environment with changing needs and requirements Ability to work independently and in a team, take initiative, and communicate effectively ADDITIONAL REQUIREMENTS: Able to work full time, onsite for a minimum of 12 consecutive weeks beginning in May or June 2027 COMPENSATION AND BENEFITS: Software Engineering Intern/Freshman/Sophomore: $30 USD per hour Software Engineering Intern/Junior/Senior: $34 USD per hour Software Engineering Intern/Completed Bachelor's: $36 USD per hour Software Engineering Intern/Completed Master's: $38 USD per hour Software Engineering Intern/Completed PhD: $40 USD per hour Your salary will be determined by academic level. Hourly pay is just one part of our total rewards package at SpaceXAI. You may also be eligible for a stipend to subsidize relocation costs, as well as access to our comprehensive medical coverage and a 401(k) retirement plan. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .

View more...
Software Engineering
VerifiedToday

Site Reliability Engineer - Memphis

On-sitefull timeSeniorMississippi, United States
Apply Now

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and facility boundaries. We are looking for candidates from power plants, nuclear power plants, data centers, or people who currently work or have worked in facilities like SpaceXAI with power, compute, cooling, and related plant systems. RESPONSIBILITIES: Own monitoring architecture and signal quality: what we alert on, suppress, and trust. Consume NOC noise-disposition feedback to drive suppression and redesign. Treat alert noise as a design failure, not an operator failure. Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene. Run blameless postmortems and drive corrective actions to closed, not filed. Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries. Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current. Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects). Define error budgets and availability objectives at campus and service boundaries as adopted by the business. Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus. PROFILE SIGNALS: Experience in power plants, nuclear power plants, data centers, or facilities like SpaceXAI — current or prior — spanning power, compute, cooling, or related plant systems. Proven large-scale incident command experience and calm technical leadership on a bridge. Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality. Treats alert noise as a design failure, not an operator failure. Cross-team facilitation and systems engineering depth across software and facility boundaries. SUCCESS MEASURED BY: MTTD / MTTR trend for SEV-class events % of SEVs with a blameless postmortem and closed action items Alert actionable ratio Recurrence rate of incident classes Monitoring coverage of critical dependencies across compute, network, storage, power, and cooling BASIC QUALIFICATIONS: Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience). 5+ years of experience in site reliability, systems engineering, plant operations, or large-scale production operations, preferably in power plants, nuclear power plants, high-performance computing, or data center environments. Proven large-scale incident command experience and calm technical leadership on a bridge. Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality. Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry. Experience writing and operating playbooks or runbooks with a 24/7 operations, control room, or NOC partner. Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them. Excellent problem-solving skills with a data-driven approach to reliability engineering. Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering. PREFERRED SKILLS AND EXPERIENCE: Experience in power plants, nuclear power plants, or other high-consequence industrial control environments. Current or prior work in data centers or facilities like SpaceXAI spanning power, compute, cooling, and related plant systems. Experience in AI/ML infrastructure or supercomputing environments. Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries. Experience running game days, dependency mapping, and closed-loop corrective action programs. Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry. Prior work in a fast-paced startup or tech company like SpaceXAI. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .

View more...
Cloud, DevOps & SRE
VerifiedToday

OT Systems Engineer (Physical Infrastructure)

On-sitefull timeMid-LevelMississippi, United States
Apply Now

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: SpaceXAI is looking for a highly skilled and versatile OT (Operational Technology) Systems Engineer to implement and support the backend infrastructure for our next-generation controls and industrial software platforms for our physical infrastructure. This role is embedded within the mission IT organization to provide support to dedicated controls engineering organizations while leveraging adjacent IT expertise, tooling, and technologies. This role will balance support of legacy controls support systems with modernization and continuous improvement initiatives. The ideal candidate thrives in high stakes settings, brings a strong sense of urgency balanced with operational excellence, and combines deep OT expertise with IT technologies such as virtualization, VDI, GitOps, network level redundancy and edge compute solutions to create streamlined and secure ICS environments. RESPONSIBILITIES: Design, deploy, and augment next-generation OT environments that support plant and launch systems across multiple sites. Install, configure, maintain, and support industry standard controls software platforms as well as internally developed HMIs. Integrate established and emergent IT technologies to simplify management, improve security, and create scalable plant environments. Deploy and maintain development, test, and staging environments to enable controlled, systematic change introduction. Provide direct support during launch, test, and production campaigns. Perform systems/software upgrades and maintenance between critical operations (including evenings and weekends as needed). Proactively monitor services and respond rapidly to incidents to maintain high availability and performance. Leverage automation tools and contribute to infrastructure-as-code and broader DevOps initiatives with a focus on ICS environments. Work with internal controls engineers to iterate on simulation and emulation environment to enhance test coverage and augment change control processes. Write and maintain standards, architectures, best practices, and documentation (including system overviews, design drawings, and operational procedures), with emphasis on highly reliable and secure industrial environments, especially at data boundaries. Collaborate with cross-functional teams (IT, security, controls, and operations) as well as external customers and integrators to design robust OT architectures and resolve technical issues. Ensure OT systems are configured and maintained in compliance with industry and cybersecurity standards (e.g., Purdue Model, IEC 62443). BASIC QUALIFICATIONS: 3+ years of experience in OT systems engineering or industrial control systems administration. Hands-on experience with multiple industry standard controls software platforms and tools. Significant experience designing, deploying, supporting, and troubleshooting OT environments in high-reliability settings. PREFERRED SKILLS AND EXPERIENCE: Experience supporting real-time systems, industrial control networks, or operational technology (OT) environments in aerospace, defense, energy, or similar high-reliability industries. Experience designing architectures that incorporate hyperconverged, rugged industrial edge, and distributed compute technologies. Working knowledge of industrial protocols, controls networks, and OT cybersecurity best practices. Proficiency in scripting (Bash/PowerShell/Python) and automation frameworks (Puppet, Terraform, Ansible, etc.). Experience with configuration management, provisioning, infrastructure as code, and DevOps concepts/tools. Familiarity with Active Directory, multi-platform authentication, and identity environments in OT contexts. System administration experience managing Windows and Linux servers, rudimentary database administration exposure, and storage/backup experience. Network administration experience and understanding of the OSI model, especially Layer 1/2/3 considerations as they apply to industrial and controls networks. Excellent communication skills with the ability to communicate with internal/external customers, vendors, and management in both formal and informal situations. ADDITIONAL REQUIREMENTS: Willingness to participate in an after-hours on-call rotation and work extended hours or weekends as necessary. Willingness to travel (up to 20%, including flying to other sites worldwide) and potential time at sea supporting marine systems. Ability to lift 30 lbs. Ability to work at heights. Ability to drive (active valid driver's license). SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .

View more...
Cloud, DevOps & SRE
VerifiedToday

OT Systems Engineer (Physical Infrastructure)

On-sitefull timeMid-LevelTennessee, United States
Apply Now

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: SpaceXAI is looking for a highly skilled and versatile OT (Operational Technology) Systems Engineer to implement and support the backend infrastructure for our next-generation controls and industrial software platforms for our physical infrastructure. This role is embedded within the mission IT organization to provide support to dedicated controls engineering organizations while leveraging adjacent IT expertise, tooling, and technologies. This role will balance support of legacy controls support systems with modernization and continuous improvement initiatives. The ideal candidate thrives in high stakes settings, brings a strong sense of urgency balanced with operational excellence, and combines deep OT expertise with IT technologies such as virtualization, VDI, GitOps, network level redundancy and edge compute solutions to create streamlined and secure ICS environments. RESPONSIBILITIES: Design, deploy, and augment next-generation OT environments that support plant and launch systems across multiple sites. Install, configure, maintain, and support industry standard controls software platforms as well as internally developed HMIs. Integrate established and emergent IT technologies to simplify management, improve security, and create scalable plant environments. Deploy and maintain development, test, and staging environments to enable controlled, systematic change introduction. Provide direct support during launch, test, and production campaigns. Perform systems/software upgrades and maintenance between critical operations (including evenings and weekends as needed). Proactively monitor services and respond rapidly to incidents to maintain high availability and performance. Leverage automation tools and contribute to infrastructure-as-code and broader DevOps initiatives with a focus on ICS environments. Work with internal controls engineers to iterate on simulation and emulation environment to enhance test coverage and augment change control processes. Write and maintain standards, architectures, best practices, and documentation (including system overviews, design drawings, and operational procedures), with emphasis on highly reliable and secure industrial environments, especially at data boundaries. Collaborate with cross-functional teams (IT, security, controls, and operations) as well as external customers and integrators to design robust OT architectures and resolve technical issues. Ensure OT systems are configured and maintained in compliance with industry and cybersecurity standards (e.g., Purdue Model, IEC 62443). BASIC QUALIFICATIONS: 3+ years of experience in OT systems engineering or industrial control systems administration. Hands-on experience with multiple industry standard controls software platforms and tools. Significant experience designing, deploying, supporting, and troubleshooting OT environments in high-reliability settings. PREFERRED SKILLS AND EXPERIENCE: Experience supporting real-time systems, industrial control networks, or operational technology (OT) environments in aerospace, defense, energy, or similar high-reliability industries. Experience designing architectures that incorporate hyperconverged, rugged industrial edge, and distributed compute technologies. Working knowledge of industrial protocols, controls networks, and OT cybersecurity best practices. Proficiency in scripting (Bash/PowerShell/Python) and automation frameworks (Puppet, Terraform, Ansible, etc.). Experience with configuration management, provisioning, infrastructure as code, and DevOps concepts/tools. Familiarity with Active Directory, multi-platform authentication, and identity environments in OT contexts. System administration experience managing Windows and Linux servers, rudimentary database administration exposure, and storage/backup experience. Network administration experience and understanding of the OSI model, especially Layer 1/2/3 considerations as they apply to industrial and controls networks. Excellent communication skills with the ability to communicate with internal/external customers, vendors, and management in both formal and informal situations. ADDITIONAL REQUIREMENTS: Willingness to participate in an after-hours on-call rotation and work extended hours or weekends as necessary. Willingness to travel (up to 20%, including flying to other sites worldwide) and potential time at sea supporting marine systems. Ability to lift 30 lbs. Ability to work at heights. Ability to drive (active valid driver's license). SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .

View more...
Cloud, DevOps & SRE
VerifiedToday

Site Reliability Engineer - Memphis

On-sitefull timeSeniorTennessee, United States
Apply Now

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and facility boundaries. We are looking for candidates from power plants, nuclear power plants, data centers, or people who currently work or have worked in facilities like SpaceXAI with power, compute, cooling, and related plant systems. RESPONSIBILITIES: Own monitoring architecture and signal quality: what we alert on, suppress, and trust. Consume NOC noise-disposition feedback to drive suppression and redesign. Treat alert noise as a design failure, not an operator failure. Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene. Run blameless postmortems and drive corrective actions to closed, not filed. Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries. Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current. Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects). Define error budgets and availability objectives at campus and service boundaries as adopted by the business. Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus. PROFILE SIGNALS: Experience in power plants, nuclear power plants, data centers, or facilities like SpaceXAI — current or prior — spanning power, compute, cooling, or related plant systems. Proven large-scale incident command experience and calm technical leadership on a bridge. Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality. Treats alert noise as a design failure, not an operator failure. Cross-team facilitation and systems engineering depth across software and facility boundaries. SUCCESS MEASURED BY: MTTD / MTTR trend for SEV-class events % of SEVs with a blameless postmortem and closed action items Alert actionable ratio Recurrence rate of incident classes Monitoring coverage of critical dependencies across compute, network, storage, power, and cooling BASIC QUALIFICATIONS: Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience). 5+ years of experience in site reliability, systems engineering, plant operations, or large-scale production operations, preferably in power plants, nuclear power plants, high-performance computing, or data center environments. Proven large-scale incident command experience and calm technical leadership on a bridge. Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality. Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry. Experience writing and operating playbooks or runbooks with a 24/7 operations, control room, or NOC partner. Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them. Excellent problem-solving skills with a data-driven approach to reliability engineering. Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering. PREFERRED SKILLS AND EXPERIENCE: Experience in power plants, nuclear power plants, or other high-consequence industrial control environments. Current or prior work in data centers or facilities like SpaceXAI spanning power, compute, cooling, and related plant systems. Experience in AI/ML infrastructure or supercomputing environments. Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries. Experience running game days, dependency mapping, and closed-loop corrective action programs. Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry. Prior work in a fast-paced startup or tech company like SpaceXAI. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .

View more...
Cloud, DevOps & SRE
VerifiedToday

Site Reliability Engineer, Data Center - Memphis

On-sitefull timeMid-LevelTennessee, United States
Apply Now

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: As a Site Reliability Engineer focused on Hardware, you will serve as an expert focused on firmware, hardware specifications, vendor relations, and failure analysis. You will proactively identify and resolve hardware issues, manage RMA processes, and stay ahead of emerging hardware technologies to support SpaceXAI's data center operations. This role demands deep technical expertise in hardware diagnostics, and forward-looking hardware evaluation. RESPONSIBILITIES: Analyze firmware packages and hardware specifications for upcoming releases for compatibility, performance, and reliability in SpaceXAI's data center environment. Run security scanning and CVE / vulnerability analysis on firmware and related components. Flag safety issues (electrical, thermal, power-protection, fail-safe behavior) before the package hits the floor. Investigate and diagnose hardware failures, including "grey failures" (ambiguous or intermittent issues), proving them as true hardware defects through rigorous testing and data analysis. Manage vendor relationships, including initiating RMA (Return Merchandise Authorization) claims, negotiating beyond standard processes when necessary, and holding vendors accountable for resolutions. Collaborate with Data Center Operations Technicians to troubleshoot, repair, and optimize hardware systems in real-time. Develop and implement monitoring tools, scripts, and processes to detect hardware anomalies early and minimize downtime. Document failure modes, RCAs, AFR / reliability models, RMA outcomes, and hardware evaluations into a team knowledge base. Participate in on-call rotations and incident response for hardware-related issues in the Memphis data center BASIC QUALIFICATIONS: Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience). 2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments. Proven expertise in firmware analysis, hardware specifications review, and release validation. Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols. Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software. Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies. Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar). Not required to be expert in all of them. Excellent problem-solving skills with a data-driven approach to reliability engineering. Ability to work collaboratively with cross-functional teams, including operations technicians. PREFERRED SKILLS AND EXPERIENCE: Experience in AI/ML infrastructure or supercomputing environments. Knowledge of vendor ecosystems (e.g., NVIDIA, Dell, HP, Supermicro) and supply chain management. Certifications in hardware engineering or reliability (e.g., CRE, CompTIA Server+). Prior work in a fast-paced startup or tech company like SpaceXAI. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .

View more...
Cloud, DevOps & SRE
VerifiedToday

Software Engineer, Data Center - Memphis

On-sitefull timeMid-LevelMississippi, United States
Apply Now

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The Data Center Engineering team builds the internal systems and platforms that keep our data centers running at the scale and reliability required for frontier AI training and inference. We partner closely with datacenter operations, research, and infrastructure teams to deliver high-leverage tools that turn raw operational data into clear insight and action. RESPONSIBILITIES: Build and operate the software stacks that make site operations scalable, auditable, and fast — including repair trackers, vendor turnback workflows, operational dashboards, and their integrations. Your users are the technicians, managers, NOC operators, and leadership who run the fleet. Design, build, and operate multi-service production systems (UI, APIs, data pipelines, auth, and on-call) for systems such as SRT-style repair/maintenance trackers and vendor turnover/turnback state machines. Own correctness of operational state: node state accuracy, queue ownership, and audit trails — the data that decides what work happens on the floor. Build and maintain integrations with ticketing, inventory/rack systems, telemetry stores, and vendor portals. Keep the tools themselves reliable: uptime, data integrity, reconciliation, access control, and safe deploys. Embed with SiteOps and NOC users; measure workflow adoption, not just feature delivery. BASIC QUALIFICATIONS: Bachelor’s degree in Computer Science, Engineering, or related fields. 3+ years building and operating production software (backend and/or full-stack). Strong fundamental knowledge of computer science - data structures, algorithms, operating systems and networking. Strong proficiency in at least one programming language e.g. Rust, Python, JavaScript, Java, C++, etc. Experience designing and developing RESTful APIs for mission-critical applications. Experience working with at least one database like Postgres, MongoDB, MySQL, DynamoDB, etc. Experience collaborating with cross-functional teams. Experience working in cloud platforms like GCP, Microsoft Azure, AWS, OCI or similar. Experience using observability tools and dashboards like New Relic, Splunk, Grafana, etc. Experience writing unit tests and integration tests. PREFERRED SKILLS AND EXPERIENCE: Full-stack or backend + data experience, especially with workflow and state-machine systems. Experience automating deployments using CI/CD tools like Azure Devops, ArgoCD, Jenkins, GitHub Actions, etc. Experience building high-correctness operational UIs where a wrong value can dispatch a human to the wrong rack. On-call discipline and a track record of treating internal platforms with production rigor. Experience designing and shipping multi-service systems (APIs, data stores, and at least one of: UI, pipelines, or auth). Experience delivering high-quality internal tools in rapidly changing environments (as a tech lead, former founder, etc.). Proven ownership of correctness-sensitive systems — state machines, workflows, or operational data where inaccurate state has real-world impact. Experience integrating with external systems via APIs (e.g. ticketing, inventory, telemetry, or vendor portals). MS in Computer Science or related field. Experience collaborating closely with operations, NOC, and infrastructure teams. Experience working with performance load testing tools like BlazeMeter, k6, etc. ADDITIONAL REQUIREMENTS: The role is fully onsite in Memphis, TN or Southhaven, MS. Candidates are expected to be located near the area or open to relocation. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .

View more...
Software Engineering
VerifiedToday

Site Reliability Engineer, Data Center - Memphis

On-sitefull timeMid-LevelMississippi, United States
Apply Now

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: As a Site Reliability Engineer focused on Hardware, you will serve as an expert focused on firmware, hardware specifications, vendor relations, and failure analysis. You will proactively identify and resolve hardware issues, manage RMA processes, and stay ahead of emerging hardware technologies to support SpaceXAI's data center operations. This role demands deep technical expertise in hardware diagnostics, and forward-looking hardware evaluation. RESPONSIBILITIES: Analyze firmware packages and hardware specifications for upcoming releases for compatibility, performance, and reliability in SpaceXAI's data center environment. Run security scanning and CVE / vulnerability analysis on firmware and related components. Flag safety issues (electrical, thermal, power-protection, fail-safe behavior) before the package hits the floor. Investigate and diagnose hardware failures, including "grey failures" (ambiguous or intermittent issues), proving them as true hardware defects through rigorous testing and data analysis. Manage vendor relationships, including initiating RMA (Return Merchandise Authorization) claims, negotiating beyond standard processes when necessary, and holding vendors accountable for resolutions. Collaborate with Data Center Operations Technicians to troubleshoot, repair, and optimize hardware systems in real-time. Develop and implement monitoring tools, scripts, and processes to detect hardware anomalies early and minimize downtime. Document failure modes, RCAs, AFR / reliability models, RMA outcomes, and hardware evaluations into a team knowledge base. Participate in on-call rotations and incident response for hardware-related issues in the Memphis data center BASIC QUALIFICATIONS: Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience). 2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments. Proven expertise in firmware analysis, hardware specifications review, and release validation. Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols. Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software. Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies. Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar). Not required to be expert in all of them. Excellent problem-solving skills with a data-driven approach to reliability engineering. Ability to work collaboratively with cross-functional teams, including operations technicians. PREFERRED SKILLS AND EXPERIENCE: Experience in AI/ML infrastructure or supercomputing environments. Knowledge of vendor ecosystems (e.g., NVIDIA, Dell, HP, Supermicro) and supply chain management. Certifications in hardware engineering or reliability (e.g., CRE, CompTIA Server+). Prior work in a fast-paced startup or tech company like SpaceXAI. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .

View more...
Cloud, DevOps & SRE
VerifiedToday

Software Engineer, Data Center - Memphis

On-sitefull timeMid-LevelTennessee, United States
Apply Now

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: The Data Center Engineering team builds the internal systems and platforms that keep our data centers running at the scale and reliability required for frontier AI training and inference. We partner closely with datacenter operations, research, and infrastructure teams to deliver high-leverage tools that turn raw operational data into clear insight and action. RESPONSIBILITIES: Build and operate the software stacks that make site operations scalable, auditable, and fast — including repair trackers, vendor turnback workflows, operational dashboards, and their integrations. Your users are the technicians, managers, NOC operators, and leadership who run the fleet. Design, build, and operate multi-service production systems (UI, APIs, data pipelines, auth, and on-call) for systems such as SRT-style repair/maintenance trackers and vendor turnover/turnback state machines. Own correctness of operational state: node state accuracy, queue ownership, and audit trails — the data that decides what work happens on the floor. Build and maintain integrations with ticketing, inventory/rack systems, telemetry stores, and vendor portals. Keep the tools themselves reliable: uptime, data integrity, reconciliation, access control, and safe deploys. Embed with SiteOps and NOC users; measure workflow adoption, not just feature delivery. BASIC QUALIFICATIONS: Bachelor’s degree in Computer Science, Engineering, or related fields. 3+ years building and operating production software (backend and/or full-stack). Strong fundamental knowledge of computer science - data structures, algorithms, operating systems and networking. Strong proficiency in at least one programming language e.g. Rust, Python, JavaScript, Java, C++, etc. Experience designing and developing RESTful APIs for mission-critical applications. Experience working with at least one database like Postgres, MongoDB, MySQL, DynamoDB, etc. Experience collaborating with cross-functional teams. Experience working in cloud platforms like GCP, Microsoft Azure, AWS, OCI or similar. Experience using observability tools and dashboards like New Relic, Splunk, Grafana, etc. Experience writing unit tests and integration tests. PREFERRED SKILLS AND EXPERIENCE: Full-stack or backend + data experience, especially with workflow and state-machine systems. Experience automating deployments using CI/CD tools like Azure Devops, ArgoCD, Jenkins, GitHub Actions, etc. Experience building high-correctness operational UIs where a wrong value can dispatch a human to the wrong rack. On-call discipline and a track record of treating internal platforms with production rigor. Experience designing and shipping multi-service systems (APIs, data stores, and at least one of: UI, pipelines, or auth). Experience delivering high-quality internal tools in rapidly changing environments (as a tech lead, former founder, etc.). Proven ownership of correctness-sensitive systems — state machines, workflows, or operational data where inaccurate state has real-world impact. Experience integrating with external systems via APIs (e.g. ticketing, inventory, telemetry, or vendor portals). MS in Computer Science or related field. Experience collaborating closely with operations, NOC, and infrastructure teams. Experience working with performance load testing tools like BlazeMeter, k6, etc. ADDITIONAL REQUIREMENTS: The role is fully onsite in Memphis, TN or Southhaven, MS. Candidates are expected to be located near the area or open to relocation. SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .

View more...
Software Engineering
VerifiedToday

Page 1 of 8

PreviousNext