401,286open jobs
13,975companies
77,769added this week
Browse all
Location
Remote (Argentina, Brazil, Colombia)
Employment
Full-Time
Overview
Company
Impact
Profile match
Solvd is a software engineering services company headquartered in Foster City, California, and founded in 2011. The company provides custom software development, quality engineering, test automation, data engineering, and AI integration services to product companies. It runs delivery teams across Eastern Europe and Latin America and serves clients in retail, financial services, media, and technology.

Solvd Inc. is a rapidly growing AI-native consulting and technology services firm delivering enterprise transformation across cloud, data, software engineering, and artificial intelligence. We work with industry-leading organizations to design, build, and operationalize technology solutions that drive measurable business outcomes.

Following the acquisition of Tooploox, a premier AI and product development company, Solvd now offers true end-to-end delivery -from strategic advisory and solution design to custom AI development and enterprise-scale implementation. Our capability centers combine deep technical expertise, proven delivery methodologies, and sector-specific knowledge to address complex business challenges quickly and effectively.

We are looking for a Production Support Engineer to join a small, high-trust team responsible for the health and reliability of a revenue-critical sales platform. You'll sit between end users - partners and call centers - and engineering, triaging incidents, managing communication, and keeping stakeholders informed and calm when things go wrong.

You'll own the incident management response process end-to-end - from first alert through post-mortem and corrective action follow-up. This is not a developer role. The technical bar is deliberately calibrated: you need to understand how APIs work, read logs, and interpret what you're seeing - not implement fixes. What matters equally is your ability to translate technical issues into plain language and manage expectations across very different audiences.

Longevity and genuine interest in the role matter here. This team values people who want to grow with it, not move through it.

What you'll do

  • Monitor platform health and triage incoming incidents - distinguishing critical issues (outages, service degradation) from non-critical ones (bugs, defects).

  • Investigate incidents using logging tools - reading API calls, responses, and log data to understand what happened and where.

  • Own the incident management response process, post-mortems, and corrective action follow-up.

  • Notify stakeholders of critical issues proactively - specifying SLA risk and communicating clearly on status via email, phone, or ticket system.

  • Cross-reference tickets across multiple systems and follow defects through the full lifecycle until closure.

  • Communicate clearly with partners and call center teams - translating technical findings into plain language and managing expectations throughout resolution.

  • Manage the incident queue in Jira and prioritize bugs within engineering sprint cycles.

  • Participate in weekly cross-functional meetings with engineering and account/call center management.

  • Provide suggestions for continual improvement of applications and processes.

  • Join on-call rotations after ramp-up - responding to alerts via OpsGenie within defined SLA windows.

Basic qualifications

  • 2+ years of troubleshooting and resolving issues for applications, servers, or infrastructure environments.

  • 2+ years of providing clear status updates on tasks, issues, and resolutions to stakeholders at multiple levels.

  • Working knowledge of how APIs function - able to read and interpret API calls and responses; experience with Postman or similar API testing tools.

  • Ability to navigate logging and observability tools such as Splunk, Datadog, or Sumo Logic.

  • Experience with SQL queries for troubleshooting and ad hoc reporting.

  • Basic comfort reading HTML and JSON, and using browser developer tools for investigation.

  • Ability to participate in technical bridge calls and follow incidents through to resolution.

  • Exceptional communication skills - able to code-switch between technical and non-technical audiences fluidly; this is the hardest skill to train and the most important one for this role.

  • Strong time management, prioritization, and organizational skills under pressure.

  • Customer service mindset - genuine interest in supporting end users and resolving issues, not just closing tickets.

  • Empathy, humility, and comfort with ambiguity - able to investigate complex issues without a clear playbook.

  • Available during U.S. Eastern business hours (9 AM - 6 PM ET); Eastern timezone strongly preferred for onboarding and on-call coordination.

  • Bachelor's degree in a related field or equivalent work experience.

Preferred qualifications

  • Experience with AWS - Cloud Practitioner level or above.

  • Familiarity with Git in a team environment.

  • Understanding of infrastructure-as-code concepts - Terraform or similar.

  • Familiarity with OpsGenie or similar alerting platforms.

  • Experience using AI tooling to amplify troubleshooting and investigation workflows.

  • Understanding of engineering deployment lifecycle and release processes.

  • Experience with on-call rotation structures and incident severity frameworks.

  • Prior exposure to partner or call center communication management during live incidents.

  • Experience in a travel, hospitality, or high-volume transactional platform environment.

What to expect when you join

  • Comprehensive onboarding documentation and a structured 6-month ramp to full self-sufficiency.

  • On-call rotations begin only when you're ready, with manager backup during early rotations.

  • Active alert window is 8 AM-1 AM Eastern; overnight suppression windows are built in.

  • SEV-1 incidents are rare - roughly once per quarter or less; the majority of the work happens during business hours.

When you join Solvd, you'll…

  • Shape real-world AI-driven projects across key industries, working with clients from startup innovation to enterprise transformation.

  • Be part of a global team with equal opportunities for collaboration across continents and cultures.

  • Thrive in an inclusive environment that prioritizes continuous learning, innovation, and ethical AI standards.

Ready to make an impact?

If you're excited to build things that matter, champion responsible AI, and grow with some of the industry’s sharpest minds. Apply today and let’s innovate together.

Solvd is an equal opportunity employer.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
401,286 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$21k – $46k per year (Estimated) • In office • Contractor • 3+ years exp • Moscow
DevOps
Git
Management
Confluence
Jira
YouTrack
QA
Postman
SoapUI
Swagger
Apply
$19k – $46k per year (Estimated) • Remote • Moscow
SQL
Design
Figma
Management
Confluence
Draw.io
Jira
Trello
QA
Postman
Apply
$88k – $201k per year (Estimated) • Remote • Contractor • 3+ years exp • Toronto
PHP
PHP
WordPress
Analytics
Tableau
Design
Figma
Sketch
Zeplin
Management
Asana
Jira
Marketing
LinkedIn
Apply
Equity • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Bangkok
Management
Jira
SharePoint
Apply
In office • Full-Time • 3+ years exp • Bachelor's Degree • Tokyo
DevOps
Incident Management
Management
ServiceNow
Apply
Senior Accountant 2 days ago
Remote • Full-Time • 5+ years exp • Bachelor's Degree
Management
QuickBooks
Apply
$40k – $114k per year (Estimated) • Remote • Full-Time • 8+ years exp
Databases
Databricks
Snowflake
AI/ML
AWS Bedrock
Gemini
LLM
Prompt Engineering
AI Agents
Anthropic
Feature Store
Function Calling
OpenAI
RAG
DevOps
AWS
Azure
CI/CD
GCP
Vector
Apply
$170k – $200k per year • Remote • Full-Time • 8+ years exp
Databases
Amazon Aurora
Amazon DocumentDB
Amazon Neptune
Apache Kafka
Cassandra
DynamoDB
MySQL
OpenSearch
Redis
DevOps
AWS
Datadog
Terraform
Amazon CloudWatch
IAM
Apply
$55k – $138k per year (Estimated) • Remote • Full-Time • 7+ years exp
Python
TypeScript
Databases
Databricks
OpenSearch
AI/ML
AWS Bedrock
DSPy
LLM
MLFlow
RAG
Semantic Search
Semantic Search
DevOps
AWS
AWS CDK
CI/CD
CloudFormation
Docker
Platform Engineering
Vector
Amazon ECS
Apply
Remote/Hybrid • Full-Time • 2+ years exp
Cybersecurity
ISO 27001
SOC 2
Apply
See all jobs
This is one of many
401,286 more open roles from verified company boards, updated every day.