377,351open jobs
9,828companies
48,306added this week
Browse all
Location
Remote/Hybrid
Seniority
Middle · 4+ years exp
Overview
Company
Impact
Profile match
Cognium is an AI-powered wealth management platform that automates portfolio monitoring, risk analysis, recommendations, compliance checks, and advisor workflows to help wealth managers scale efficiently.

We are looking for an AI Engineer who can move fluidly between research and production: someone who can prototype a new model or prompting approach quickly, then harden it into a reliable, observable, cost-efficient service that runs at scale. You'll work closely with product and full-stack engineering to turn ambiguous problems into shipped AI features.

Responsibilities:

  • Design, build, and deploy machine learning and LLM-based systems, including retrieval-augmented generation (RAG) pipelines, fine-tuned models, and agentic workflows.
  • Own the full lifecycle of AI features: data collection and evaluation, model/prompt selection, experimentation, deployment, and monitoring in production.
  • Build and maintain data pipelines, embeddings stores, and vector databases that support search, retrieval, and personalisation.
  • Evaluate and integrate third-party model APIs (OpenAI, Anthropic, etc. ) alongside open-source and self-hosted models, balancing quality, latency, and cost.
  • Fine-tune, train, and deploy open-source models on our own infrastructure, including data prep, training/fine-tuning runs, quantisation, and self-hosted inference serving at scale.
  • Establish evaluation frameworks and metrics to measure model quality, drift, hallucination rate, and business impact over time.
  • Collaborate with full-stack engineers to expose AI capabilities through clean, well-documented APIs and SDKs.
  • Implement guardrails, safety checks, and monitoring for AI systems running in production.
  • Stay current with the fast-moving AI/ML landscape and bring back practical recommendations on tools, models, and techniques.
  • Write clear technical documentation and communicate trade-offs to both technical and non-technical stakeholders.

Requirements:

  • 4+ years of experience building and shipping machine learning or AI-powered systems in production.
  • Strong Python skills, with hands-on experience using frameworks such as PyTorch, TensorFlow, or similar.
  • Practical experience with large language models prompting, fine-tuning, RAG, or agent frameworks (e. g., LangChain, LlamaIndex, or custom implementations).
  • Hands-on experience training and fine-tuning open-source models (e. g., Llama, Mistral, Qwen) full fine-tuning or parameter-efficient methods (LoRA/QLoRA) and deploying them on your own infrastructure rather than relying solely on hosted APIs.
  • Experience standing up self-hosted inference serving on your own infra (e. g., vLLM, TGI, Triton, Ray Serve), including GPU provisioning, batching, and cost/latency optimisation.
  • Solid understanding of ML fundamentals: model evaluation, overfitting, data leakage, and experimentation methodology.
  • Experience with vector databases and embedding-based retrieval (e. g., Pinecone, Weaviate, pgvector, FAISS).
  • Comfort working with cloud infrastructure (AWS, GCP, or Azure) and containerised deployments (Docker, Kubernetes), including GPU-backed compute.
  • Familiarity with MLOps practices: model versioning, CI/CD for ML, monitoring, and rollback strategies for both hosted and self-hosted models.
  • Strong software engineering fundamentals: clean code, testing, code review, and API design.
  • Proficiency working with AI coding assistants (e. g., Claude Code, GitHub Copilot, Cursor) to accelerate development, while critically reviewing and validating AI-generated code.
  • Excellent communication skills and comfort working in a remote, async-friendly environment.

Nice to Have:

  • Experience with model quantisation, distillation, or optimisation techniques (e. g., GPTQ, AWQ, ONNX) for efficient self-hosted inference.
  • Background in NLP, computer vision, or recommendation systems.
  • Experience with streaming data pipelines (Kafka, Spark) or feature stores.
  • Contributions to open-source ML/AI tooling.
  • Prior experience in a startup or fast-growth environment.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
377,351 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
AI Engineer 2 hours ago
Remote/Hybrid • 4+ years exp
Python
Databases
Apache Kafka
FAISS
pgvector
Pinecone
Weaviate
PostgreSQL
AI/ML
AI Agents
Anthropic
AWQ
Claude
Claude Code
Computer Vision
Copilot
Cursor
Embeddings
Feature Store
Fine-tuning
GPTQ
Hallucination
Knowledge Distillation
LangChain
Llama
LlamaIndex
LLM
LLM Guardrails
LoRA
Mistral
NLP
ONNX
OpenAI
PyTorch
QLoRA
Qwen
RAG
Ray
Ray Serve
Recommender Systems
Spark
TensorFlow
TGI
Triton
vLLM
PEFT
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Kubernetes
Vector
Apply
$26k – $64k per year (Estimated) • In office • Full-Time • 8+ years exp • PhD • Hyderabad
Java
Python
Python
Boto3
AI/ML
Agentforce
AI Agents
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GCP
Hyper-V
Kubernetes
KVM
Platform Engineering
Self-Healing
Terraform
VMWare
Analytics
Tableau
Marketing
Salesforce
Apply
Lead AI Engineer 2 hours ago
$36k – $86k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
Python
FastAPI
Databases
Databricks
AI/ML
AI Agents
Hallucination
LangChain
LangGraph
LLM
Multimodal AI
PyTorch
RAG
Semantic Kernel
Transformers
DevOps
AWS
Azure
Vector
Apply
$107k – $257k per year (Estimated) • In office • Full-Time • Sydney • Brisbane • Melbourne
Java
Python
Scala
SQL
Databases
Snowflake
AI/ML
dbt
DevOps
AWS
Azure
CI/CD
Cortex
Datadog
GCP
Splunk
Terraform
Vector
Prometheus
Cybersecurity
GDPR
HIPAA
SOC 2
Analytics
ETL/ELT
Apply
$31k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Hyderabad
Python
SQL
Databases
Google BigQuery
Snowflake
AI/ML
AI Agents
LLM
Prompt Engineering
RAG
DevOps
Azure
CI/CD
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
377,351 more open roles from verified company boards, updated every day.