Open roles/Engineering/Staff Software Engineer, AI Inference Gateway
Live

Staff Software Engineer, AI Inference Gateway

Compute Exchange — the world's first open exchange for compute, connecting buyers who need GPU capacity with the suppliers who have it · 10 employees · $5M raised

Remote — US, North America, or Europe (US time zone overlap required)Full-time · Remote$200K – $275K base + meaningful early-stage equity

About the role

Compute Exchange is building the world's first open exchange for compute — a marketplace connecting buyers who need GPU capacity with the suppliers who have it, routing across roughly 100 GPU providers. 10 employees, $5M raised.

This is a greenfield opportunity to architect and build a new AI inference gateway product from the ground up — high-ownership, and ideal for a former founder or entrepreneurial engineer who wants to take a product from proof-of-concept to market.

Remote, working US time zone hours (a 5-hour overlap window to accommodate a team across North America and Europe). Open to visa transfers (e.g. OPT, H1B transfers).

What you will own

  • Architect the inference gateway end-to-end — routing, load balancing, and failover across ~100 GPU providers
  • Build routing logic that leverages proprietary cost and performance data
  • Own the metering, usage accounting, and enterprise billing layer
  • Establish the observability, reliability, and multi-tenancy foundations
  • Drive the product from proof-of-concept to market

Requirements

Must haveArchitected and shipped products to market at high-growth startups
Must have7–20 years of experience in backend/infrastructure engineering, building distributed systems (Go, Rust, or Python)
Must haveStrong software architecture and system design skills for distributed systems
Must haveProficient in Go, Rust, or a similar systems language, and comfortable with Python
Must haveKnowledge of running inference workloads — inference servers such as vLLM, SGLang, TensorRT-LLM, or TGI — and an understanding of what actually drives GPU utilization and latency
Must haveKubernetes and cloud infrastructure experience (autoscaling, multi-region)
Nice to haveExperience with marketplaces, exchanges, or other two-sided systems where pricing and supply are dynamic
Nice to haveBS in Computer Science or equivalent engineering discipline
Nice to haveExperience building or operating inference gateways, usage-based metering, or token routing systems
Not a fitPrimarily a data scientist or ML engineer without systems engineering depth
Not a fitOnly large-company or enterprise experience with no startup or founder-type work

We interview honestly against this list — if you meet most of the “Must have” rows, apply.

Compensation

Remote — US, North America, or Europe (US time zone overlap required)
Location
Full-time · Remote
Employment
$200K – $275K base + meaningful early-stage equity
Compensation
Competitive equity
Equity

Role details

LocationRemote — US, North America, or Europe (US time zone overlap required)
EmploymentFull-time · Remote
Compensation$200K – $275K base + meaningful early-stage equity
EquityCompetitive equity
DepartmentEngineering
VisaOpen to visa transfers (e.g. OPT, H1B transfers)
Vikki

Vikki

Your recruiter for this role

Not sure you tick every box? Reach out before you apply — happy to tell you what actually matters here.