401,286open jobs
13,975companies
77,769added this week
Browse all
Location
Remote/Hybrid (South Africa)
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Machine Learning Engineer - AI Architecture Research based in South Africa.

This role focuses on researching and building next-generation AI model architectures that can move from experimental concepts to scalable production systems.

You will work at the intersection of machine learning research, model engineering, and real-world deployment.

The position offers the opportunity to challenge established architectural assumptions and explore alternatives to conventional Transformer-based designs.

You will design experiments, prototype new neural networks, and evaluate trade-offs across compute, memory, latency, and model performance.

The role involves close collaboration with inference and systems engineers to make research ideas efficient and deployable.

You will also contribute to research reproduction, benchmarking, technical exploration, and potentially open-source work.

This is an opportunity to have meaningful influence on AI architecture while working in a fast-moving, research-oriented environment.

Accountabilities:

    • Research and develop novel neural network architectures, including alternatives or extensions to Transformers, recurrent and hybrid models, and long-context systems.
    • Design and execute architecture-level experiments focused on scaling laws, memory mechanisms, training behavior, and compute-performance trade-offs.
    • Prototype models end-to-end, translating research concepts into robust, training-ready implementations.
    • Analyze model behavior, failure modes, inductive biases, and architectural strengths and limitations.
    • Collaborate with inference and systems engineering teams to ensure new architectures are efficient, scalable, and suitable for deployment.
    • Read, reproduce, evaluate, and extend cutting-edge machine learning research papers.
    • Contribute to internal research notes, benchmarks, experiments, and open-source initiatives where applicable.
    • Move fluidly between theoretical investigation, rapid experimentation, and production-oriented engineering.
    • Requirements:

      • Strong foundation in machine learning and deep learning fundamentals, with practical experience applying them to model development.
      • Hands-on experience implementing neural network or model architectures from scratch.
      • Strong understanding of attention mechanisms, RNNs, state-space models, hybrid architectures, or related approaches.
      • Solid knowledge of training dynamics, optimization, scaling behavior, and architecture-level performance considerations.
      • Understanding of model-level memory, latency, compute, and efficiency constraints.
      • Proficiency with PyTorch or JAX and the ability to develop and experiment with research-oriented ML code.
      • Ability to evaluate architectural ideas through both theoretical reasoning and empirical experimentation.
      • Strong communication skills, with the ability to clearly explain technical concepts and architectural trade-offs.
      • Preferred experience with non-Transformer architectures such as RNN variants, state-space models, or long-context systems.
      • Preferred background in research-driven startups, open-source machine learning projects, large-scale training, or custom training loops.
      • Publications, preprints, notable research contributions, or experience with inference optimization and deployment constraints are advantageous.
      • Benefits:

        • Competitive compensation and meaningful equity.
        • Opportunity to work directly on core AI model architecture rather than focusing primarily on fine-tuning.
        • Significant influence over technical and research direction within a rapidly growing organization.
        • Small, high-caliber team with fast feedback loops and a strong research-oriented environment.
        • Opportunity to take research concepts from experimentation through to production deployment.
        • Full-time position with a globally distributed work environment.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
401,286 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$130k – $180k per year • Remote • 10+ years exp • Bachelor's Degree
C++
C
C++
LLVM
PyTorch C++
TensorFlow C++
C
MPI
AI/ML
CUDA
CUDA Toolkit
CUTLASS
DeepSpeed
JAX
MLIR
NCCL
PyTorch
ROCm
TensorFlow
TensorRT
Triton
vLLM
DevOps
AWS
Azure
GCP
HPC
Apply
$130k – $180k per year • Remote • 10+ years exp • Master's Degree
Python
AI/ML
AI Agents
DeepSpeed
DPO
Fine-tuning
FSDP
Hallucination
Knowledge Graph
LangChain
LangGraph
LlamaIndex
LLM
LoRA
Multimodal AI
NLP
PEFT
PPO
PyTorch
QLoRA
RAG
Ray
RLHF
SFT
Synthetic Data
Transformers
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Vector
Apply
$130k – $180k per year • Remote • 10+ years exp • Bachelor's Degree
Python
AI/ML
AI Agents
Gymnasium
JAX
LLM
ML-Agents
OpenAI
PyTorch
Ray
Reinforcement Learning
Reward Modeling
RLHF
RLlib
Stable-Baselines3
TensorFlow
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Robotics
Imitation Learning
Isaac Sim
MuJoCo
NVIDIA Omniverse
Reinforcement Learning
Apply
$130k – $180k per year • Remote • 10+ years exp • Bachelor's Degree
Python
Databases
FAISS
Milvus
Pinecone
Weaviate
AI/ML
AI Agents
Computer Vision
Embeddings
Fine-tuning
JAX
LangChain
LangGraph
LlamaIndex
LLM
Multimodal AI
NLP
Prompt Engineering
PyTorch
RAG
Ray
Recommender Systems
TensorFlow
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Apply
$80k – $100k per year • Remote • 6+ years exp • Master's Degree
Python
AI/ML
AI Agents
Fine-tuning
JAX
Multimodal AI
PyTorch
RAG
Apply
Remote • Full-Time • 10+ years exp
Apply
Remote • Full-Time • 10+ years exp
Apply
$28k – $71k per year (Estimated) • Remote • Contractor • 5+ years exp • Bachelor's Degree
Python
AI/ML
Amazon SageMaker
Evidently AI
MLFlow
PyTorch
Recommender Systems
TensorFlow
DevOps
AWS
CI/CD
GitLab
GitLab CI
Grafana
Prometheus
Apply
$50k – $127k per year (Estimated) • Remote • Contractor • 5+ years exp • Bachelor's Degree
Python
AI/ML
Amazon SageMaker
Evidently AI
MLFlow
PyTorch
Recommender Systems
TensorFlow
DevOps
AWS
CI/CD
GitLab
GitLab CI
Grafana
Prometheus
Apply
Remote/Hybrid • Internship
Apply
See all jobs
This is one of many
401,286 more open roles from verified company boards, updated every day.