One of the most common mistakes VPS administrators make is finding out about server problems only after they have already caused damage. CPU pegged at 100% slowing your application to a crawl, RAM exhausted causing processes to get killed, or a full disk preventing your database from writing — all of these are preventable with proper monitoring alerts. The key difference between reacting to outages and preventing them is a well-configured alert system that notifies you before things escalate to a crisis.
This guide walks through setting up an alert system on a Linux VPS from scratch: choosing the right tool, configuring thresholds, sending notifications via Telegram and Email, and testing that everything works. The examples cover CPU, Memory, Disk, and service availability — applicable to any VPS running WordPress, Node.js, Docker, or business applications.
Monitoring vs Alerting: Why You Need Both
Many VPS administrators set up a beautiful Grafana dashboard with real-time CPU and memory graphs but find it useless when a problem occurs at 3 AM because nobody is watching the screen. A good alert system acts as your eyes on the server around the clock.
- Monitoring — Collects metrics such as CPU%, RAM usage, and I/O throughput continuously and stores historical data for trend analysis.
- Alerting — Compares metrics against defined thresholds and dispatches notifications through your chosen channel when limits are crossed.
- Together — Monitoring tells you what happened; alerting tells you what needs attention right now.
A practical example: Your disk reaches 90% full at 8 PM. Without alerts you might not notice until morning. With an alert, your phone buzzes with a Telegram message immediately. You SSH in, clear old log files, and disk drops to 60% before the database ever stops writing. Zero downtime, zero data loss.
Choosing the Right Tool for VPS Alerting
Several solid options exist for Linux VPS monitoring and alerting. The best choice depends on how many servers you manage and how much flexibility you need.
| Tool | Complexity | Best For | Notifications |
|---|---|---|---|
| Netdata | Easy | Single VPS, getting started | Email, Telegram, Slack, PagerDuty |
| Prometheus + Alertmanager | Medium–High | Multiple servers, custom app metrics | Email, Telegram, Webhook, PagerDuty |
| Zabbix | High | Enterprise, large infrastructure | Email, SMS, Webhook |
| Shell Script + Cron | Easy (DIY) | Targeted alerts without extra software | Email, Telegram, Line Notify |
For a typical Ubuntu/Debian VPS, start with Netdata — one command installs a complete monitoring and alerting stack with a built-in dashboard. Graduate to Prometheus when you need full control or manage multiple servers.
Installing Netdata and Configuring Telegram Alerts
Netdata is the fastest path from zero to a working alert system. It collects over a thousand metrics by default, displays a real-time dashboard, and fires alerts without any external database or additional components.
Install Netdata
wget -O /tmp/netdata-kickstart.sh https://my-netdata.io/kickstart.sh sudo bash /tmp/netdata-kickstart.sh --stable-channel --disable-telemetry
After installation, Netdata runs at http://your-vps-ip:19999. Restrict access with a firewall rule so it is not publicly accessible.
Configure Telegram Notifications
Edit /etc/netdata/health_alarm_notify.conf:
SEND_TELEGRAM="YES" TELEGRAM_BOT_TOKEN="YOUR_BOT_TOKEN_HERE" TELEGRAM_CHAT_ID="YOUR_CHAT_ID_HERE" DEFAULT_RECIPIENT_TELEGRAM="YOUR_CHAT_ID_HERE"
Customize CPU Alert Thresholds
Alert rules live in /etc/netdata/health.d/. Create a custom CPU rule:
# /etc/netdata/health.d/cpu-custom.conf alarm: cpu_usage_warning on: system.cpu lookup: average -15m unaligned of user,system,softirq,irq,guest units: % every: 1m warn: $this > 80 crit: $this > 90 delay: down 15m multiplier 1.5 max 1h info: CPU usage exceeds defined threshold to: sysadmin
Reload without restarting:
sudo kill -USR2 $(pidof netdata)
Prometheus + Alertmanager for Advanced Alerting
If you need granular control over alert routing, inhibition, silencing, and grouping, or you manage multiple servers, Prometheus with Alertmanager is the industry standard. The setup involves more steps but gives you far more flexibility.
Install Node Exporter to Collect System Metrics
wget https://github.com/prometheus/node_exporter/releases/download/v1.7.0/node_exporter-1.7.0.linux-amd64.tar.gz tar xf node_exporter-1.7.0.linux-amd64.tar.gz sudo mv node_exporter-1.7.0.linux-amd64/node_exporter /usr/local/bin/ sudo tee /etc/systemd/system/node_exporter.service <<'EOF' [Unit] Description=Node Exporter After=network.target [Service] User=node_exporter ExecStart=/usr/local/bin/node_exporter Restart=on-failure [Install] WantedBy=multi-user.target EOF sudo useradd -rs /bin/false node_exporter sudo systemctl enable --now node_exporter
Alert Rules for Memory and Disk
# /etc/prometheus/rules/server_alerts.yml
groups:
- name: server_alerts
rules:
- alert: HighMemoryUsage
expr: (1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)) * 100 > 85
for: 5m
labels:
severity: warning
annotations:
summary: "High memory on {{ $labels.instance }}"
description: "Memory used: {{ printf \"%.1f\" $value }}%"
- alert: DiskSpaceLow
expr: (1 - (node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"} / node_filesystem_size_bytes{fstype!~"tmpfs|overlay"})) * 100 > 80
for: 5m
labels:
severity: warning
annotations:
summary: "Low disk space on {{ $labels.instance }}"
description: "{{ $labels.mountpoint }} is {{ printf \"%.1f\" $value }}% full"
- alert: InstanceDown
expr: up == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Instance {{ $labels.instance }} is down"
description: "{{ $labels.instance }} has been unreachable for 1 minute"
Configuring Alertmanager for Telegram Delivery
With alert rules defined in Prometheus, configure Alertmanager to route notifications to Telegram:
# /etc/alertmanager/alertmanager.yml
global:
resolve_timeout: 5m
route:
group_by: ['alertname', 'instance']
group_wait: 30s
group_interval: 5m
repeat_interval: 3h
receiver: 'telegram-notifications'
receivers:
- name: 'telegram-notifications'
telegram_configs:
- bot_token: 'YOUR_BOT_TOKEN'
chat_id: YOUR_CHAT_ID
parse_mode: 'HTML'
message: |
<b>{{ .Status | toUpper }}</b> - {{ .GroupLabels.alertname }}
{{ range .Alerts }}
<b>Instance:</b> {{ .Labels.instance }}
<b>Details:</b> {{ .Annotations.description }}
{{ end }}
inhibit_rules:
- source_match:
severity: 'critical'
target_match:
severity: 'warning'
equal: ['alertname', 'instance']
The inhibit_rules section prevents warning-level spam when a critical alert is already active for the same instance — a key feature for reducing alert fatigue in production environments.
Minimal Shell Script Alerting with Cron
If you prefer not to install any additional software, a simple shell script combined with cron provides effective alerting for the most critical thresholds. This approach is fully self-contained and easy to understand.
#!/bin/bash
# /opt/monitor/check_resources.sh
TELEGRAM_TOKEN="YOUR_BOT_TOKEN"
CHAT_ID="YOUR_CHAT_ID"
HOSTNAME=$(hostname)
send_alert() {
curl -s -X POST "https://api.telegram.org/bot${TELEGRAM_TOKEN}/sendMessage" \
-d chat_id="${CHAT_ID}" \
-d parse_mode="HTML" \
-d text="$1" >/dev/null
}
# CPU check (5-minute load average vs core count)
CPU_LOAD=$(awk '{print $1}' /proc/loadavg)
CPU_CORES=$(nproc)
CPU_PERCENT=$(echo "scale=0; $CPU_LOAD * 100 / $CPU_CORES" | bc)
[ "$CPU_PERCENT" -gt 90 ] && send_alert "🔴 <b>[CRITICAL] CPU</b> - ${HOSTNAME}: ${CPU_PERCENT}%"
[ "$CPU_PERCENT" -gt 80 ] && [ "$CPU_PERCENT" -le 90 ] && \
send_alert "🟡 <b>[WARNING] CPU</b> - ${HOSTNAME}: ${CPU_PERCENT}%"
# Memory check
MEM_TOTAL=$(grep MemTotal /proc/meminfo | awk '{print $2}')
MEM_AVAIL=$(grep MemAvailable /proc/meminfo | awk '{print $2}')
MEM_PCT=$(echo "scale=0; (($MEM_TOTAL - $MEM_AVAIL) * 100) / $MEM_TOTAL" | bc)
[ "$MEM_PCT" -gt 90 ] && send_alert "🔴 <b>[CRITICAL] Memory</b> - ${HOSTNAME}: ${MEM_PCT}%"
[ "$MEM_PCT" -gt 80 ] && [ "$MEM_PCT" -le 90 ] && \
send_alert "🟡 <b>[WARNING] Memory</b> - ${HOSTNAME}: ${MEM_PCT}%"
# Disk check
DISK_PCT=$(df / | awk 'NR==2{print $5}' | sed 's/%//')
[ "$DISK_PCT" -gt 90 ] && send_alert "🔴 <b>[CRITICAL] Disk</b> - ${HOSTNAME}: ${DISK_PCT}%"
[ "$DISK_PCT" -gt 80 ] && [ "$DISK_PCT" -le 90 ] && \
send_alert "🟡 <b>[WARNING] Disk</b> - ${HOSTNAME}: ${DISK_PCT}%"
Schedule with cron for checks every 5 minutes:
chmod +x /opt/monitor/check_resources.sh crontab -e # Add: */5 * * * * /opt/monitor/check_resources.sh
Prevent Alert Fatigue: If CPU stays high for hours, a script running every 5 minutes will flood you with dozens of messages. Add a lock file with a cooldown period: if [ ! -f /tmp/cpu_alert ] || [ $(( $(date +%s) - $(stat -c %Y /tmp/cpu_alert) )) -gt 3600 ]; then send_alert "..."; touch /tmp/cpu_alert; fi — this ensures the same alert fires no more than once per hour until the condition clears.
Essential Alert Checklist for Every Production VPS
Regardless of which tool you choose, these alerts should be active on every VPS serving real traffic:
System Resource Alerts
- CPU Usage — 15-minute average above 80% (Warning), above 90% (Critical)
- Memory Usage — RAM utilization above 85% (Warning), above 95% (Critical)
- Disk Space — Root or data partition above 80% (Warning), above 90% (Critical)
- Disk I/O Wait — Sustained high I/O wait can indicate disk failure approaching
- Swap Usage — Swap above 50% means RAM is insufficient; consider upgrading
Service Availability Alerts
- Web Server Down — Nginx or Apache not responding (HTTP 5xx or connection refused)
- Database Unreachable — MySQL or PostgreSQL not accepting connections
- Application Process Missing — Node.js or PHP-FPM processes no longer running
- SSL Certificate Expiry — Certificate expiring within 14 days
Security Alerts
- Multiple Failed SSH Logins — Brute force attack in progress
- Successful Root Login — Anyone logging in as root successfully from an unknown IP
- Unusual Outbound Traffic — Possible compromise or cryptominer running
Testing Your Alert System Before Relying on It
A misconfigured alert system that silently fails is worse than having no alerts at all — it creates a false sense of security. Always test end-to-end before trusting the system with production workloads.
Test Telegram Connectivity
curl -s -X POST "https://api.telegram.org/bot${TELEGRAM_TOKEN}/sendMessage" \
-d chat_id="${CHAT_ID}" \
-d text="✅ Alert system test from $(hostname) - OK"
Simulate High CPU Load
sudo apt install stress-ng -y stress-ng --cpu 4 --timeout 300s & # Wait 5–15 minutes for alerts to fire based on your configured duration # Stop the test: kill %1
Check Alertmanager and Netdata Logs
# Alertmanager sudo journalctl -u alertmanager -f # Netdata health log tail -f /var/log/netdata/health.log
Frequently Asked Questions (FAQ)
What CPU threshold should I use for alerts?
For general-purpose servers, a Warning at 80% and Critical at 90% works well. Measure the average over 5–15 minutes rather than instant spikes, since brief CPU bursts are normal. If you are receiving too many alerts (alert fatigue), try increasing the sustained duration to 10 minutes before adjusting thresholds. Fine-tune based on your actual workload patterns over time.
What is the difference between Netdata and Prometheus Alertmanager?
Netdata is ideal for a single VPS that needs a real-time dashboard and alerting in one tool — install with one command and it works immediately. Prometheus plus Alertmanager suits multi-server environments or situations requiring custom application metrics. It offers more flexibility but requires more configuration. For a standalone VPS just getting started, Netdata provides better value out of the box.
Why are my alerts not being sent even when CPU is very high?
Common causes include a misconfigured notification channel (wrong Telegram bot token or chat ID), alert rule syntax errors that prevent Alertmanager from loading the rules file, outbound internet blocked by a firewall rule on port 443, or alerts being in an Inhibit or Silence state unintentionally. Check Alertmanager logs and test with amtool alert add before relying on the system in production.
How do I send VPS alerts through a Telegram bot?
Create a bot via @BotFather on Telegram to receive an API token, then add the bot to the group or chat where you want notifications delivered. Retrieve the Chat ID by calling https://api.telegram.org/bot<TOKEN>/getUpdates after sending a message. Configure Alertmanager receivers or Netdata notification settings with the token and chat_id values. Always verify with a test curl request before deploying to production.
High-Performance KVM VPS by AsiaGB
AsiaGB VPS includes full root access, supports Docker, MySQL, Python, and Node.js. Plans starting from 500 THB/month.
View VPS Plans