We're hiring a hands-on Computer Vision and Multimodal AI Scientist to join our applied AI team, working under our principal applied scientist on production CV and multimodal systems. This is an individual-contributor role for someone who wants to go deep on hard technical problems, not manage people or run a business. You'll own specific workstreams end-to-end: literature review, architecture design, and production deployment. You'll work on problems spanning classical CV (detection, tracking, pose estimation), modern multimodal/LMM systems (vision-language models, GUI/screen understanding, contextual captioning), and increasingly, agentic AI models that perceive, reason, and act.
Responsibilities:
- Design and ship production CV/multimodal pipelines for object detection, tracking, pose estimation, image stitching, and face/material/scene understanding for real client use cases, not just PoCs.
- Build and fine-tune vision-language and multimodal models (VLMs/LMMs) for tasks like contextual image description, document/OCR understanding, and screen/GUI comprehension.
- Contribute to agentic AI systems where visual perception feeds downstream reasoning and action (e. g., computer-use agents, autonomous GUI interaction).
- Own the synthetic data and annotation strategy where labeled data is scarce or expensive, including GAN-based synthesis and semi-automated annotation pipelines.
- Take models from ~60-70% baseline accuracy to production-grade (95%+ mAP/accuracy) through targeted data augmentation, architecture iteration, and error analysis, not just hyperparameter tuning.
- Work alongside client stakeholders (up to director/VP level) to translate ambiguous business problems into scoped technical solutions, under guidance from the team lead.
- Collaborate closely with the principal applied scientist on architecture decisions and with other engineers/annotators on execution.
- Stay current with the field vision transformers, self-supervised learning, and LoRA/fine-tuning techniques, and bring relevant advances into active projects.
Requirements:
- 8+ years of hands-on experience in computer vision, with direct exposure to at least 3 of: object detection/tracking, pose estimation, image stitching/registration, face recognition/matching, document AI/OCR, or material/scene classification.
- Proven track record of shipping CV models to production, not just research or academic benchmarks. Be ready to discuss the gap between your best validation metric and what actually held up in production and how you closed it.
- Strong practical fluency in PyTorch/TensorFlow, modern CV architectures (CNNs, Vision Transformers, and CLIP/DINO-style backbones), and classical CV (OpenCV).
- Working experience with multimodal AI / LMMs, VLM fine-tuning, RAG for vision-language tasks, or multimodal retrieval.
- Experience with synthetic data generation (GANs, simulation, or augmentation pipelines) to solve data-scarcity problems.
- Solid software engineering fundamentals in Python at a minimum, comfortable working in Git/Docker/Linux environments, and writing code that other engineers can build on.
- Demonstrated ability to work independently on ambiguous problems and communicate technical tradeoffs to non-technical stakeholders.
Nice to Have (Strong Signal, Not Gatekeeping):
- Exposure to agentic AI/LLM agents like LangChain, LangGraph, tool-use/function-calling, or computer-use agent architectures.
- Patents or peer-reviewed publications (CVPR/ICCV/ECCV/ICME/IJCAI tier or equivalent industry venues).
- Experience in regulated domains: HIPAA/GDPR-compliant systems, especially in healthcare or banking.
- Experience with on-device/edge deployment (mobile NDK, LiteRT/MediaPipe, embedded inference).
- Experience working in a fast-paced, resource-constrained environment (startup or lean team), even if not in a founding role.

