About the Role
As a Senior Software Engineer within our Commodities tribe, you will lead the end-to-end integration of US inland waterway barge data into Kpler’s global cargo intelligence platform. Operating under a "you build it, you run it" philosophy, you will design robust ingestion paths, stream real-time data using Kafka, and connect barge movement models directly to our customer-facing surfaces. This role carries genuine design authority, empowering you to solve complex entity resolution problems while mentoring peers and elevating technical craft across our crews.
Key Responsibilities
Lead end-to-end integration architecture: Drive the full engineering lifecycle-from external feed ingestion and domain modeling to streaming, entity resolution, and UI distribution.
Design robust ingestion & entity-resolution systems: Map incoming barge data onto Kpler reference data (vessels, zones, installations, products) using existing Elasticsearch and PostgreSQL matching patterns.
Build scalable Kafka streaming pipelines: Construct event-driven contracts, schemas (Avro), deduplication mechanisms, and dead-letter queue replays with Python and Scala.
Deliver data across customer surfaces: Surface integrated barge data cleanly across Elasticsearch read models, external APIs, data warehouse feeds, and the Kpler Terminal (TypeScript/Vue).
Champion production quality & observability: Maintain full operational ownership of your services, setting SLOs, participating in on-call rotations, and driving post-incident improvements.
Ensure data integrity & analyst alignment: Validate external feeds and integrate human-in-the-loop analyst workflows so corrections seamlessly apply to barge movements.
Mentor and elevate engineering craft: Guide and mentor crew members, document architectural designs, and actively contribute to decomposing legacy estates into event-driven services.
Experience & Background
Production data engineering experience: ~5+ years of experience designing, operating, and maintaining data-intensive production systems with full ownership.
Advanced Python & SQL expertise: Proven ability to build clean, well-bounded services within large, complex, and evolving codebases.
Event-driven streaming proficiency: Hands-on experience with Kafka streaming in production, including schema management (Avro/Schema Registry), deduplication, and replay patterns.
Third-party integration & entity resolution: Demonstrated background integrating external data feeds, managing schema mapping, and building data quality monitoring frameworks.
Operational mindset: Direct experience managing production observability, writing documentation, and resolving incidents in a distributed system environment.
Functional programming & Scala: Experience or familiarity with Scala (Kafka Streams) or a track record of rapidly mastering new languages.
Geospatial & domain knowledge: Familiarity with geospatial data tools (PostGIS) or exposure to maritime, logistics, or commodities domains.
Modern data stack exposure: Experience with Elasticsearch, Airflow, Snowflake/lakehouse table formats, Kubernetes, and GitOps deployments.
UI/Frontend collaboration: Working knowledge of TypeScript or Vue to collaborate on user-facing terminal features.

