HPC Maintenance and Technical Support — Managed Operations
Enterprise HPC cluster maintenance, monitoring, software updates and 24/7 technical support. Proactive monitoring that acts before a failure surfaces.
High performance computing infrastructure requires continuous maintenance and expert operations once it is in production. Hardware failures, software incompatibilities, performance regressions and security vulnerabilities — handling all of that calls for a full-time HPC systems specialist. Mevasis HPC maintenance services take that burden on entirely.
Scope of Service
Proactive Monitoring
Prometheus + Grafana: CPU, memory, storage and network metrics
DCGM Exporter: GPU health and performance (power, temperature, error counts)
SLURM Exporter: queue depth, pending jobs, resource utilisation
node_exporter: hardware sensors, disk SMART status
Alertmanager: automatic notification on threshold breach (email, SMS, Slack)
When thresholds are crossed — CPU utilisation sustained above 95%, or disk health degrading, for example — the Mevasis team begins work before the customer is aware of it.
Software Updates and Patch Management
- Operating system security patches, applied in coordinated maintenance windows
- SLURM, MPI libraries and Lmod updates
- CUDA and driver updates
- Update support for application software (GROMACS, OpenFOAM and similar)
Hardware Maintenance
- Disk failure detection and replacement
- Memory error tracking (ECC log analysis)
- GPU health assessment
- Network switch and InfiniBand port status tracking
- Planned maintenance (fans, filters, thermal paste)
User and Workload Management
- SLURM account and partition management
- Job queue priority policies
- Usage reporting by department and project
- Detection and handling of problem jobs
Support Levels
| Level | Definition | Response Time |
|---|---|---|
| Critical | System completely unreachable | 4 hours |
| High | Significant component failure with user impact | Within the business day |
| Medium | Performance degradation, partial impact | 2 business days |
| Low | Configuration request, improvement | 5 business days |
Monthly Reporting
A detailed operations report is delivered at the end of every month:
- Uptime and SLA compliance
- Queue depth and pending job statistics
- Resource usage trends (CPU/GPU/storage)
- Summary of maintenance performed
- Recommended capacity or configuration improvements
To learn more about Mevasis HPC maintenance services, request a quote or contact our team for a free assessment of your current infrastructure.
Frequently Asked Questions
Proactive monitoring, software updates and patching, hardware fault diagnosis, capacity and performance reporting, user management support, and technical intervention within the SLA.
Both models exist. Most interventions are carried out remotely; for hardware failures and critical situations a team is dispatched to site.
Four hours for critical faults (system down), within the business day for high-priority issues, and within one business day for standard requests.
Related Solutions
Let's Build Future Together.
