MLOps/DevOps Engineer

Melio is seeking a passionate DevOps Engineer to join our expanding team:

  • Level: Intermediate or Senior
  • Position: Full time
  • Salary: Based on technical experience
  • Location: South Africa (Remote)

Melio AI is an AI product and consulting company, recognised as an AWS Advanced Tier Partner specialising in AI/ML. We drive real-world AI adoption for businesses across industries. We don’t just build models - we deliver production-ready AI solutions that work at scale. From AI Agents and LLM-powered applications to high-performance MLOps, we develop AI systems that deliver tangible business value.

We’re looking for a MLOps/DevOps Engineer to join our growing Cloud & Infrastructure team. If you’re someone who thrives on autonomy, wants to lead infrastructure projects, and is passionate about AWS, cloud-native technologies, and site reliability, this is your chance to make an impact. Beyond client projects, you’ll also be contributing to our internal AI marketplace, Highwind, where we build, deploy, and scale AI-driven solutions.

Job Description

You will be working with clients on projects focused on delivering reliable application software to production efficiently with a cloud-native approach. This includes managing and implementing infrastructure and deployment processes for new and existing projects. This requires the candidate to work closely with developers, QA, and operational teams to ensure robust infrastructure and site reliability.

Melio believes in nurturing cross-functional capabilities in our team, so you will need to work closely with other technical roles, such as software engineers, data scientists, machine learning engineers, and other DevOps engineers either to observe or to assist.

This is a technical role, but due to the consulting nature of many of our projects, the candidate needs to be able to communicate effectively with both business and technical stakeholders.

We work extensively within the AWS cloud ecosystem and help clients architect, deploy, and manage AWS-native solutions. The candidate will be exposed to evaluating and integrating various AWS services with clients’ existing technology stacks, focusing on leveraging the full capabilities of AWS infrastructure.

What you will be working on

  • AWS Cloud Infrastructure – Design, implement, and manage scalable AWS infrastructure for production applications.
  • MLOps – Build and operate ML deployment pipelines, model serving infrastructure, monitoring, and scaling for production AI systems.
  • AgentOps – Support the deployment, observability, and reliability of AI agents and LLM-powered applications in production.
  • CI/CD & Deployment Automation – Build and maintain pipelines for reliable, automated software delivery across all environments.
  • Cloud-Native Architecture – Deploy and manage containerised workloads using Kubernetes and other CNCF technologies.
  • Site Reliability & Monitoring – Implement observability, alerting, and incident response practices to keep systems running smoothly.
  • Infrastructure as Code – Automate infrastructure provisioning and configuration using modern IaC tooling, such as AWS CDK, Terraform and CloudFormation.
  • Deploying solutions into our AI Marketplace – Support the infrastructure and deployment of AI solutions on our Highwind marketplace, a platform enabling businesses to access plug-and-play AI APIs.
  • Cross-Functional Collaboration – Work with developers, data scientists, and business teams to deliver reliable production systems.

What we are looking for

  • Passion for Cloud & DevOps – You stay current with AWS and cloud-native trends, love automating infrastructure, and care about reliability at scale.
  • Software Engineering Mindset – You see infrastructure as code—you care about maintainability, repeatability, and best practices.
  • Autonomous & Leadership-Driven – You take initiative, own your work, and are eager to grow into leading infrastructure projects.
  • Collaborative Yet Independent – You can drive solutions independently but also work closely with teammates to solve complex challenges.

What you would be assisting other team members with (secondary responsibilities)

  • Assist with backend development from time to time.
  • Help data scientists and machine learning engineers deploy their models into production.
  • Maintain day-to-day management and administration of projects.

Qualification & Experience

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Software Engineering, or related field.
  • 3+ years experience in a DevOps/SRE role, or related computer engineering field.
  • 2+ years software development experience (e.g. Python, Go, Java, or JavaScript) and proficiency with Bash.
  • 2+ years hands-on AWS experience (e.g. EC2, ECS/EKS, Lambda, S3, IAM).
  • Experience maintaining production environments, including monitoring, incident response, and operational reliability.
  • Experience with Infrastructure as Code (AWS CDK, Terraform, or CloudFormation), CI/CD pipelines, Docker, and container orchestration.
  • 1+ years experience with Kubernetes or other CNCF projects.
  • Experience (or strong interest) in MLOps and AgentOps: deploying, monitoring, and operating ML models and AI agents in production.

Nice to Have

  • Open-source contributions, startup, side project, or product development experience.
  • AWS certifications (e.g. Solutions Architect, SysOps Administrator) or CKA/CKAD.
  • Hands-on experience with AWS CDK (TypeScript or Python).
  • Experience with AWS Bedrock and Bedrock AgentCore.

Personal Attributes

  • Up-to-date on latest industry trends; able to articulate trends clearly and confidently.
  • Able to interact with other team members via code and design documents.
  • Good interpersonal skills and communication with all levels of management.
  • Able to multitask, prioritize, and manage time efficiently.
  • Curious and eager to learn about new technologies.
  • Strong in critical thinking and problem-solving.

What you will learn and grow into

  • Scaling Cloud Infrastructure – Learn how to take infrastructure from prototype to full-scale production on AWS.
  • Enterprise Cloud Deployments – Gain experience working with real-world systems for top companies.
  • End-to-End Infrastructure Ownership – Move beyond day-to-day ops to architecting cloud solutions.
  • Architecting Cloud-Native Solutions – Design scalable architectures that integrate with production systems and enterprise workflows.
  • Leading Infrastructure Project Delivery – Take ownership of projects, from scoping requirements to delivering solutions that drive measurable impact.

Why Join Us?

  • Work with one of South Africa’s leading AI startups.
  • Build infrastructure that ships—not just proof-of-concept environments.
  • Contribute to Highwind, our AI marketplace, where your work scales beyond individual projects.
  • Take full ownership of your projects and grow into a cloud infrastructure leader.

Contact Us

If you are interested in this position please email Harry (harry@melio.ai) with the below information:

  • CV
  • Expected Salary
  • Notice period