Nebula is the system of state for enterprise agent workflows. This role owns the storage and execution systems beneath it: object-storage-native architecture, distributed execution, partitioning, multi-region deployment, and customer-managed infrastructure. It's a product engineering role, not general DevOps or internal-platform work.
What you'll work on
- Evolve Nebula's S3-backed storage architecture, including logs, manifests, snapshots, indexes, checkpoints, caching, and compaction.
- Build durable ingestion and workflow execution on top of queues, including idempotency, retries, backpressure, replay, dependency scheduling, and partial-failure recovery.
- Design partitioning and routing across tenants, workers, Kubernetes clusters, storage partitions, and regions.
- Define clear consistency, durability, and recovery guarantees across asynchronous pipelines and replicated object storage.
- Improve read-path performance and keep p95 and p99 latency predictable as datasets and workloads grow.
- Build multi-region replication, routing, failover, and recovery mechanisms.
- Detect and recover from queue buildup, retry storms, stale indexes, data drift, hotspots, noisy tenants, and silent correctness failures.
- Operate the same core data plane across our SaaS, BYOC, on-premises, and air-gapped deployments.
- Work directly with the founders and early customers to turn scale, reliability, and deployment constraints into product architecture.
What we're looking for
- Production experience building distributed systems, storage systems, databases, workflow engines, search or indexing infrastructure, or high-throughput backend systems.
- Experience owning stateful systems where correctness, latency, durability, and cost mattered simultaneously.
- Strong judgment around consistency, concurrency, persistence, caching, partitioning, replication, and failure recovery.
- Experience with queues and asynchronous execution, including retries, idempotency, ordering, replay, and backpressure.
- Hands-on experience with Kubernetes and cloud infrastructure.
- Experience debugging difficult production problems such as tail latency, queue buildup, retry storms, storage bottlenecks, data drift, hotspots, and partial failures.
- High ownership and comfort turning ambiguous systems problems into explicit invariants, guarantees, and measurable outcomes.
We care more about demonstrated systems judgment than a particular number of years of experience.
Nice to have
- Experience with Rust, Go, C++, or another systems-oriented language.
- Experience with object-storage-native databases, distributed databases, event streaming, workflow orchestration, vector indexes, graph systems, or search infrastructure.
- Experience with multi-region systems and asynchronous replication.
- Experience supporting BYOC, on-premises, air-gapped, or regulated customer environments.
- Experience building infrastructure for AI agents or other data-intensive AI applications.
Why Zeroset
- Own core systems that directly determine the reliability, performance, and economics of the product.
- Help define a new architecture for durable agent state rather than maintaining a mature system with fixed assumptions.
- Work with a small, technical team with high trust, low bureaucracy, and direct access to customers and founders.
- Take on difficult storage and distributed-systems problems at the center of production enterprise AI.