Role Description
As a member of the Grid Communications and Platform team, you will develop a control system for our distributed fleet of devices. This system will enable performant communications with our devices, which includes the ingestion of millions of events per day, operations including sensor data retrievals and device configuration updates, and the work of other teams reliant on our device's data.
You will own, test, operate, and plan the future of our services end to end. Along the way, you’ll partner closely with other Software teams such as Data Engineering and Devops, as well as cross-functional teams such as Firmware, Operations, and Data Science. As an early member of the GCAP team, your technical and nontechnical decisions will influence the long term architecture and tradeoffs of our system.
Responsibilities
- Design, build, test, and operate software that is scaleable, observable, secure, and fault
tolerant.
- Create low-latency, event-driven pipelines for high-volume device telemetry and command
processing, and make key architectural decisions from protocol design to infrastructure.
- Own our simulated and hardware in the loop testing software in collaboration with the
Firmware team.
- Collaborate with Data Science and Operations teams to ensure your designs account for a
variety of environmental, connectivity, and operating conditions.
- Lead cross functional projects to optimize system performance end to end.
- Own observability, monitoring, and incident response capabilities to support reliable
production operations.
Required Skills
- 5+ years of hands-on experience developing distributed, event-driven systems on cloud-native platforms using Python or Go.
- Experience with Kafka, Kinesis, Redpanda, or Apache Pulsar.
- Experience with using observability, monitoring, and logging tools such as Grafana, Prometheus, Loki, or similar.
- Strong communication skills and the ability to navigate ambiguity.
Bonus Skills
- Production experience optimizing transport layer protocols (TCP/UDP/QUIC).
- Deep understanding of networking, DNS, TLS, and DTLS.
- Experience developing software for distributed physical, IoT, or sensing systems.
- Experience with Apollo Router / GraphQL federation gateways.
- Expertise with Kubernetes, GitOps workflows, and Infrastructure as Code.
- Experience in high-growth startup environments where you must wear many hats.

