Senior MLOps + DevOps Engineer (On-Prem AI Platform)
REQ #0261
3 Openings
Full-Time • Direct Hire
Actively Hiring
Role Overview & Requirements
Verified Requisition
Job Title: Senior MLOps + DevOps Engineer (On-Prem AI Platform)
Role Overview:
We are looking for a Senior MLOps + DevOps Engineer (8+ years) to architect, build, and
scale AI/ML platforms in an on-prem enterprise environment.
This role requires end-to-end ownership of ML systems, infrastructure, CI/CD, and
production reliability, enabling scalable deployment of machine learning and GenAI
solutions.
Key Responsibilities:
1. Platform Architecture & Ownership
- Design and own end-to-end ML platform architecture (data → training → deployment →
monitoring)
- Define and enforce best practices for scalable and secure ML systems
- Standardize MLOps + DevOps frameworks and processes
2. Model Deployment & Serving
- Deploy and manage ML/LLM models on GPU-based on-prem infrastructure
- Optimize inference performance (latency, throughput, batching)
- Implement model versioning, A/B testing, and rollback strategies
3. CI/CD & Automation
- Design and implement CI/CD pipelines for ML models, APIs, and data workflows
- Enable automated testing, deployment, and release management
4. Infrastructure & Containerization
- Manage Linux-based (RHEL preferred) on-prem infrastructure
- Containerize applications using Docker
- Deploy and orchestrate workloads using Kubernetes / OpenShift
- Operate within restricted or air-gapped environments
5. Data & System Integration
- Build pipelines integrating structured databases and high-volume logs/streaming data
- Support batch and real-time inference architectures
6. Monitoring, Observability & Reliability
- Implement end-to-end observability (model + infra)
- Use tools like Prometheus, Grafana, ELK stack
- Ensure high availability, SLA adherence, and incident response
7. GenAI & Advanced ML Systems
- Deploy RAG pipelines and vector databases
- Manage LLM serving frameworks
- Work with agent orchestration frameworks
8. Leadership & Collaboration
- Mentor engineers on MLOps and DevOps best practices
- Collaborate with cross-functional teams
- Drive design reviews and production readiness
Required Skills:
- Strong Python and scripting (Bash)
- Deep understanding of ML lifecycle and productionization
- Experience deploying ML/LLM systems in production
- Linux, Docker, Kubernetes/OpenShift
- CI/CD tools (Jenkins/GitLab CI)
- SQL and data pipeline experience
Good to Have:
- GPU optimization knowledge
- MLflow / Kubeflow
- Terraform / Ansible
- Experience in on-prem or restricted environments
Experience:
- 8+ years in MLOps / DevOps / Platform Engineering
- Proven experience scaling production ML systems
Ideal Candidate:
A hands-on platform architect who can operate across ML systems and infrastructure,
driving automation, scalability, and reliability.
Our 3-Step Hiring Process
Step 01
Apply & AI Screen
Submit your resume and contact info for automated pipeline matching.
Step 02
Recruiter Connect
30-min conversation covering your goals, projects, and work alignment.
Step 03
Team Round & Offer
Technical deep-dive with leadership, followed by decision and offer.
Why You'll Love Working Here
Competitive Base Salary & Annual Bonuses
Comprehensive Medical & Health Coverage
Flexible Work-From-Home & Hybrid Culture
Generous Paid Time Off & Company Holidays
Continuous Upskilling & Conference Stipend
Top-tier Hardware & Home Office Setup
Related Positions
Explore similar career opportunities at Comspark Innov Infra
Discover how next-gen AI resume parsing and Kanban pipelines transform hiring.
Official Platform Demo Video
Video preview will be available here soon.
Feature 01
AI Resume Parser
Feature 02
Visual Kanban
Feature 03
Excel Importer
Feature 04
Anti-Duplicate
Spark ATS Privacy Preference Center
Manage your consent preferences for cookies and browser storage
When you use Spark ATS, we store or access information on your browser in the form of cookies. You can customize which categories you wish to permit below. Essential cookies cannot be deactivated because they are strictly required for security, session continuity, and core application operations.
Strictly Essential Cookies2 cookies
Required for core authentication, session continuity, CSRF security, and tenant isolation. These cookies cannot be turned off.
Always Active
Cookie Name
Purpose
Duration
Type
PHPSESSID
Maintains user authentication, active login state, and links requests to your tenant organization context.
Session
First-party HTTP
spark_cookie_consent
Stores your cookie consent choices (essential, functional, analytics, marketing) and policy version (1.0).
12 Months
First-party HTTP
Functional & Preference Cookies2 cookies
Enables voluntary convenience features that persist across browser restarts, such as the "Remember Me" login function.
Cookie Name
Purpose
Duration
Type
spark_ats_remember
Secure hashed authentication selector token for persistent "Remember Me" sign-in across browser sessions.
30 Days
First-party HTTP
spark_ats_remember_user
Remembers your username or email address for convenient login form prefill on your device.
30 Days
First-party HTTP
Analytics Cookies0 cookies
Used to measure aggregate portal usage and site performance. Spark ATS currently does not deploy any analytics trackers.
No analytics trackers or telemetry cookies are currently deployed by Spark ATS.
Advertising & Marketing Cookies0 cookies
Used to track visitors across websites to deliver targeted advertising. Spark ATS does not deploy advertising trackers or pixels.
No marketing, retargeting, or advertising cookies are deployed by Spark ATS.
Complete, verified inventory of all cookies currently utilized by the Spark ATS platform. Every cookie listed below is first-party and scoped directly to this application.
Cookie Name
Category
Duration
Provider
Purpose
PHPSESSID
Essential
Session
First-Party (Spark ATS)
Maintains user authentication, session security, CSRF defense, and tenant workspace routing.
spark_cookie_consent
Essential
12 Months
First-Party (Spark ATS)
Stores your category consent selections and policy versioning to prevent repetitive prompts.
spark_ats_remember
Functional
30 Days
First-Party (Spark ATS)
Hashed authentication selector token for voluntary persistent "Remember Me" login.
spark_ats_remember_user
Functional
30 Days
First-Party (Spark ATS)
Remembers username/email for automatic login form prefilling when Remember Me is enabled.
Zero Third-Party Trackers Guarantee: Spark ATS does not embed third-party analytics (e.g. Google Analytics), tracking pixels (e.g. Meta Pixel), or advertising brokers. All cookies are first-party and scoped strictly to this service domain.
Learn how Spark ATS respects your privacy, safeguards recruitment data, and transparently manages browser storage.
What are cookies?
Cookies are small text files placed on your device by websites you visit. They enable applications to remember your login session, protect forms against attacks, and store voluntary preferences.
Why Spark ATS uses cookies
To maintain secure multi-tenant isolation, authenticate recruiters and administrators, and ensure platform continuity. Optional cookies are only activated upon your explicit consent.
Essential vs. Optional
Essential technical cookies are strictly required for security and cannot be disabled. Optional cookies (Functional, Analytics, Marketing) are disabled by default and require your affirmative consent.
Managing Your Preferences
You can modify or revoke your consent choices at any time by opening this Privacy Preference Center via the "Cookie Settings" link in the footer or directly on our dedicated Cookie Policy page.