About TensorWave
Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.
About the Role
We're hiring a Network Automation Software Lead to build and own the end-to-end zero-touch provisioning (ZTP) and automation platform that stands up and operates our GPU network fabrics and to lead the small team of engineers and SREs building it.
The goal is zero human intervention and full fabric validation: a switch goes from rack-and-stack to production-ready automatically, ensuring every link across the fabric is validated against the plan and spec. Our Network Engineering team is your customer; you build the platform, they run the network on top of it.
You'll report to a Software Engineering Lead, which means this is real software engineering, not scripts bolted onto a NOC. Source control, testing, CI/CD, code review, release management, and on-call are the baseline. You'll also be responsible for making sure the platform integrates cleanly with the rest of our software and platform tooling.
What You’ll Do
- Own the end-to-end ZTP pipeline: bare-metal switch boot → image + base config → registration in source of truth → full intended config → validation → production — with zero human intervention.
- Build intent-based config generation off a network source of truth / IPAM, with GitOps-style deployment, pre/post-change validation, and safe rollout and rollback.
- Establish network validation and pre-deployment testing (snapshot/digital-twin testing) so changes are caught before they hit production fabrics.
- Build streaming telemetry and metrics/logging pipelines (gNMI / OpenConfig) for fabric health.
- Instrument what matters for GPU networks: RoCE health (PFC/ECN counters), optics and link errors, BGP / EVPN state, capacity and utilization.
- Deliver dashboards and alerting the network team actually uses — signal, not noise.
- Gather requirements, build self-service APIs and interfaces, and relentlessly accelerate their deployment velocity.
- Partner closely so the tooling reflects how the network is actually operated and turned up.
- Hire, mentor, and grow a small team of software engineers and SREs; own roadmap, prioritization, and delivery — while still carrying a meaningful share of the code yourself.
- Set technical direction and standards, and ensure clean integration points with the broader platform stack (infra provisioning, CI/CD, secrets, identity, existing observability).
- Bring software engineering rigor to network automation: code review, testing, release management, and on-call ownership.
Who You Are
Required Qualifications
- 8+ years of relevant experience
- Proven experience building network automation at scale — ideally at a hyperscaler, large cloud, or large-scale datacenter / AI-infrastructure operator.
- You've built or been a core contributor to a ZTP / device-provisioning system end-to-end, not just maintained one.
- Strong software engineering fundamentals: Python and/or Go, with real production practices (version control, testing, CI/CD, code review).
- Hands-on depth with datacenter Clos fabrics and the protocols that run them: BGP, EVPN/VXLAN, and ideally RoCEv2 / RDMA for GPU networks at scale.
- Fluency with modern network automation tech: gNMI/gNOI, OpenConfig/YANG, NETCONF; source-of-truth systems (NetBox / Nautobot); NOS platforms (SONiC/FRR or vendor equivalents); tooling like Nornir / NAPALM / Ansible.
- Experience with observability / telemetry pipelines (Prometheus, Grafana, Kafka, OpenTelemetry, or similar).
- Comfort running services on Kubernetes / containers.
- Leadership: you've led a team or been the clear technical owner of a platform, and you instinctively treat internal users as customers.
Preferred Qualifications
- Experience with GPU / AI training or inference clusters and their backend networks.
- Familiarity with the AMD networking ecosystem (Pensando DPUs, Ultra Ethernet) or building on Ethernet-based RDMA fabrics.
- Whitebox / disaggregated networking and SONiC at scale.
- Network validation / digital-twin tooling (e.g., Batfish, containerlab).
- Multi-site / multi-region datacenter buildouts.
What We Offer
- Stock Options
- 100% paid Medical, Dental, and Vision insurance for Employees
- Company Health Savings Account Contributions
- 100% paid Short Term and Long Term Disability Insurance for Employees
- Life and Voluntary Supplemental Insurance Options
- Other Insurance Options, such as Pet & Legal Insurance
- Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support
- Flexible Spending Account
- 401(k)
- Employee Assistance Program
- Flexible PTO
- Paid Holidays
- Parental Leave
- Other In-Office Perks
Equal Employment Opportunity
TensorWave is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of any protected status under applicable law.
Reasonable Accommodations
TensorWave provides reasonable accommodations in accordance with applicable laws. If you require accommodation during the hiring process, please contact [email protected].
Employment Eligibility
All offers of employment are contingent upon verification of identity and authorization to work in United States, as required by law.
Background Checks
Where permitted by law, employment may be contingent upon the successful completion of a job-related background check.
Data Privacy Notice
By submitting an application, you acknowledge that TensorWave may collect, use, and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.