Senior MLOps + DevOps Engineer (On-Prem AI Platform)

REQ #0261 3 Openings Full-Time • Direct Hire Actively Hiring

Role Overview & Requirements

Verified Requisition
Job Title: Senior MLOps + DevOps Engineer (On-Prem AI Platform) Role Overview: We are looking for a Senior MLOps + DevOps Engineer (8+ years) to architect, build, and scale AI/ML platforms in an on-prem enterprise environment. This role requires end-to-end ownership of ML systems, infrastructure, CI/CD, and production reliability, enabling scalable deployment of machine learning and GenAI solutions. Key Responsibilities: 1. Platform Architecture & Ownership - Design and own end-to-end ML platform architecture (data → training → deployment → monitoring) - Define and enforce best practices for scalable and secure ML systems - Standardize MLOps + DevOps frameworks and processes 2. Model Deployment & Serving - Deploy and manage ML/LLM models on GPU-based on-prem infrastructure - Optimize inference performance (latency, throughput, batching) - Implement model versioning, A/B testing, and rollback strategies 3. CI/CD & Automation - Design and implement CI/CD pipelines for ML models, APIs, and data workflows - Enable automated testing, deployment, and release management 4. Infrastructure & Containerization - Manage Linux-based (RHEL preferred) on-prem infrastructure - Containerize applications using Docker - Deploy and orchestrate workloads using Kubernetes / OpenShift - Operate within restricted or air-gapped environments 5. Data & System Integration - Build pipelines integrating structured databases and high-volume logs/streaming data - Support batch and real-time inference architectures 6. Monitoring, Observability & Reliability - Implement end-to-end observability (model + infra) - Use tools like Prometheus, Grafana, ELK stack - Ensure high availability, SLA adherence, and incident response 7. GenAI & Advanced ML Systems - Deploy RAG pipelines and vector databases - Manage LLM serving frameworks - Work with agent orchestration frameworks 8. Leadership & Collaboration - Mentor engineers on MLOps and DevOps best practices - Collaborate with cross-functional teams - Drive design reviews and production readiness Required Skills: - Strong Python and scripting (Bash) - Deep understanding of ML lifecycle and productionization - Experience deploying ML/LLM systems in production - Linux, Docker, Kubernetes/OpenShift - CI/CD tools (Jenkins/GitLab CI) - SQL and data pipeline experience Good to Have: - GPU optimization knowledge - MLflow / Kubeflow - Terraform / Ansible - Experience in on-prem or restricted environments Experience: - 8+ years in MLOps / DevOps / Platform Engineering - Proven experience scaling production ML systems Ideal Candidate: A hands-on platform architect who can operate across ML systems and infrastructure, driving automation, scalability, and reliability.

Our 3-Step Hiring Process

Step 01
Apply & AI Screen

Submit your resume and contact info for automated pipeline matching.

Step 02
Recruiter Connect

30-min conversation covering your goals, projects, and work alignment.

Step 03
Team Round & Offer

Technical deep-dive with leadership, followed by decision and offer.

Why You'll Love Working Here
Competitive Base Salary & Annual Bonuses
Comprehensive Medical & Health Coverage
Flexible Work-From-Home & Hybrid Culture
Generous Paid Time Off & Company Holidays
Continuous Upskilling & Conference Stipend
Top-tier Hardware & Home Office Setup

Related Positions

Explore similar career opportunities at Comspark Innov Infra
Browse All Open Positions
REQ #0264 4 Openings

Software Engineer

Comspark Innov Infra
REQ #0241 2 Openings

Biomedical Engineer

Comspark Innov Infra
REQ #0236 2 Openings

Plant & Maintenance Engineer

Comspark Innov Infra
Requisition Summary
ATS #261
Target Position
Senior MLOps + DevOps Engineer (On-Prem AI Platform)
Openings Available
3 Openings
Employment Type
Full-Time • Direct Hire
Application Status
Open • Actively Reviewing
Browse All Open Positions
100% Direct candidate submission • No middleman