About CNTXT
CNTXT is building voice AI infrastructure for the Arabic-speaking world, focusing on natural speech synthesis, real-time transcription, and conversational voice systems with high-quality Arabic language support.
Role Overview
CNTXT is seeking an AI engineer or researcher passionate about speech and voice technology. This hands-on role sits at the intersection of research and engineering and involves evaluating architectures, running fine-tuning experiments, and deploying model improvements to production.
What the Team Works On
Speech Synthesis: Build and fine-tune Arabic TTS systems using autoregressive and non-autoregressive generative architectures, neural vocoders, audio codecs, acoustic encoders, and diffusion-based audio decoders.
Speech Recognition: Develop accurate, low-latency Arabic ASR systems using encoder-decoder and CTC-based models, with streaming inference, domain adaptation, and dialect robustness.
Speech-to-Speech: Build real-time voice interaction pipelines combining ASR, language understanding, and TTS, including VAD, speaker diarization, and speech enhancement.
Arabic Language Challenges: Improve diacritization, pronunciation accuracy, dialect coverage, and model performance across MSA, Gulf, Levantine, Egyptian, and Maghrebi Arabic.
Key Responsibilities
Benchmark and evaluate Arabic TTS and ASR models using WER, speaker similarity, naturalness, and dialect-coverage metrics.
Fine-tune pretrained TTS models using curated Arabic speech data.
Run experiments involving diacritized and undiacritized input, dialect-specific training splits, and voice prompt conditioning.
Evaluate audio tokenizers, neural codecs, RVQ representations, and continuous latent approaches.
Build and maintain Arabic speech data pipelines covering audio sourcing, normalization, diacritization, filtering, and manifest generation.
Optimize TTS and ASR models for production using streaming generation, KV-cache tuning, quantization, and batched inference.
Integrate and evaluate speech-to-speech pipelines combining ASR, LLM, and TTS components.
Required Skills
Strong foundations in machine learning and deep learning.
Hands-on experience training or fine-tuning neural models.
Proficiency with Python, PyTorch, and the Hugging Face ecosystem.
Ability to read research papers and independently translate ideas into experiments.
Strong communication skills across research and engineering teams.
Nice to Have
Native or fluent Arabic language skills.
Experience with ASR, TTS, speaker verification, audio codecs, VAD, diarization, or speech enhancement.
Knowledge of Arabic linguistic structure, diacritization, and NLP preprocessing.
Experience with quantization, speculative decoding, CUDA kernels, vLLM, or TensorRT.
Publications or open-source contributions in speech or audio AI.
What CNTXT Offers
Work at the frontier of Arabic voice AI.
Direct influence on product and research direction.
A small, focused team with strong ownership and impact.
Competitive compensation and remote flexibility.

