Role Overview
Presight is seeking an experienced Senior Machine Learning Engineer with deep expertise in audio and speech technologies to design, develop, and deploy advanced machine learning models for its data integration platform.
The role focuses on building scalable, production-grade speech and audio intelligence capabilities, including speech-to-text, speaker identification, keyword spotting, language identification, and deepfake detection.
Key Responsibilities
Develop, train, and deploy production-grade machine learning models for speech-to-text, including domain adaptation.
Build speaker identification and speaker verification capabilities.
Develop keyword spotting and language identification models.
Research and implement deepfake and synthetic audio detection solutions.
Integrate transformer-based and multimodal model architectures into production pipelines.
Deploy machine learning models using NVIDIA Triton Inference Server and Kubernetes.
Optimize model serving for low latency, scalability, and real-time performance.
Design and maintain Airflow DAGs for preprocessing, feature extraction, and model enrichment workflows.
Improve models through retraining, performance monitoring, and data feedback loops.
Collaborate with Data Engineering, Backend, Platform, and Product teams.
Mentor machine learning engineers and researchers.
Oversee validation and release processes for machine learning components.
Write and review production-level Python code, with C++, Golang, or Rust considered beneficial.
Required Qualifications
Bachelor's or Master's degree in Computer Science, Machine Learning, Electrical Engineering, or a related field.
Experience leading machine learning initiatives from research through production deployment.
At least 5 years of professional Python experience.
At least 4 years of experience with PyTorch.
At least 4 years of experience with speech-to-text technologies.
Strong experience with audio processing, machine learning, Kubernetes, and Airflow.
Experience in startup or fast-paced product environments is preferred.
Preferred Skills
Deepfake and synthetic audio detection.
Speaker search, identification, and verification.
Knowledge of C++, Golang, or Rust.
Experience with NVIDIA Triton Inference Server.

