
在VPS上常见的痛点之一,就是应用在服务器重启或崩溃后无法自动恢复。systemd——每个现代Linux发行版内置的初始化系统和服务管理器——通过将您的应用作为后台服务运行来解决这个问题:开机自动启动,故障后自动重启。
什么是systemd服务单元?
服务单元是一个以 .service 为扩展名的纯文本文件,存放在 /etc/systemd/system/ 目录中。它描述了要运行哪个进程、以哪个用户身份运行、使用哪些环境变量,以及如何处理重启。systemd 读取该文件并自动管理进程的整个生命周期。
为什么不直接用PM2? PM2 在 Node.js 领域很流行,但 systemd 是无需额外安装的 Linux 原生方案。它支持任何编程语言,与 journald 集成实现集中式日志管理,是 Ubuntu 上管理服务的标准方式。
服务单元文件结构
每个 .service 文件包含三个区段:
[Unit]
Description=Human-readable service name
After=network.target # Wait for network before starting
[Service]
Type=simple # Process type
User=ubuntu # Run as this user (not root)
WorkingDirectory=/home/ubuntu/myapp
ExecStart=/usr/bin/node server.js
Restart=always # Restart whenever the process exits
RestartSec=5 # Wait 5s before restarting
[Install]
WantedBy=multi-user.target # Enable at boot
示例一:Node.js 应用
第一步:创建服务文件
sudo nano /etc/systemd/system/myapp.service
[Unit]
Description=My Node.js Application
After=network.target
[Service]
Type=simple
User=ubuntu
WorkingDirectory=/home/ubuntu/myapp
ExecStart=/usr/bin/node /home/ubuntu/myapp/server.js
Restart=on-failure
RestartSec=10
StandardOutput=journal
StandardError=journal
SyslogIdentifier=myapp
Environment=NODE_ENV=production
Environment=PORT=3000
[Install]
WantedBy=multi-user.target
第二步:启用并启动服务
# Reload systemd to pick up the new file
sudo systemctl daemon-reload
# Enable auto-start on boot
sudo systemctl enable myapp
# Start the service now
sudo systemctl start myapp
# Check status
sudo systemctl status myapp
在systemd中管理环境变量与密钥
将API密钥、数据库密码或其他敏感信息直接硬编码到单元文件中是不安全的,因为任何能执行 systemctl cat myapp 的人都能读取它们。systemd 提供了一种安全的方式,可以在不暴露到单元文件的情况下将密钥注入服务。
方案一:使用 EnvironmentFile(推荐)
# Create a separate environment file
sudo nano /etc/myapp.env
# Contents of /etc/myapp.env (no "export" keyword — KEY=VALUE only)
DB_PASSWORD=SuperSecretPassword123
API_KEY=sk-xxxxxxxxxxxxxxxxxxxx
PORT=3000
# Restrict permissions so only root can read it
sudo chmod 600 /etc/myapp.env
sudo chown root:root /etc/myapp.env
在单元文件中使用 EnvironmentFile 引用该文件:
[Service]
User=ubuntu
EnvironmentFile=/etc/myapp.env
ExecStart=/usr/bin/node /home/ubuntu/myapp/server.js
# DB_PASSWORD, API_KEY, and PORT are injected automatically
方案二:systemd 凭据(Ubuntu 22.04+)
# Encrypt a secret with systemd-creds (optionally sealed to TPM2)
echo -n "SuperSecretPassword" | sudo systemd-creds encrypt --name=db-password -
# Reference in unit file
[Service]
LoadCredential=db-password:/etc/credentials/db-password
# App reads the secret at $CREDENTIALS_DIRECTORY/db-password
安全规则:切勿将密码直接以 Environment=DB_PASS=secret 的形式写入单元文件。执行 systemctl show myapp 会暴露所有环境变量。请始终使用设有严格权限的 EnvironmentFile 代替。
示例二:Python 应用(FastAPI / Flask)
第一步:创建虚拟环境
cd /home/ubuntu/myapi
python3 -m venv venv
source venv/bin/activate
pip install fastapi uvicorn
第二步:创建Python/FastAPI服务文件
sudo nano /etc/systemd/system/myapi.service
[Unit]
Description=FastAPI Application
After=network.target
[Service]
Type=simple
User=ubuntu
WorkingDirectory=/home/ubuntu/myapi
ExecStart=/home/ubuntu/myapi/venv/bin/uvicorn main:app --host 0.0.0.0 --port 8000
Restart=always
RestartSec=5
StandardOutput=journal
StandardError=journal
SyslogIdentifier=myapi
Environment=ENV=production
[Install]
WantedBy=multi-user.target
sudo systemctl daemon-reload
sudo systemctl enable myapi
sudo systemctl start myapi
sudo systemctl status myapi
常用服务管理命令
# Check status
sudo systemctl status myapp
# Stop / Start / Restart
sudo systemctl stop myapp
sudo systemctl start myapp
sudo systemctl restart myapp
# Disable from auto-start
sudo systemctl disable myapp
# View last 100 log lines
sudo journalctl -u myapp -n 100
# Follow logs in real time
sudo journalctl -u myapp -f
生产环境最佳实践
选择合适的重启策略
Restart=always— 进程因任何原因停止时均重启Restart=on-failure— 仅在进程以非零退出码退出时重启Restart=on-abnormal— 在收到信号终止或超时时重启
资源限制
[Service]
# Cap memory at 512 MB
MemoryMax=512M
# Limit CPU to 50%
CPUQuota=50%
安全提示:始终在 [Service] 中指定专用的非root用户 User=。创建一个无登录shell的系统用户:sudo useradd -r -s /bin/false appuser,并以该用户身份运行服务以最小化权限。
重启策略对比表
选择正确的重启策略对生产环境的稳定性至关重要。重启过于激进可能掩盖应用bug;而从不重启则意味着崩溃的服务将永远无法恢复。
| 重启策略 | 触发重启的情况 | 不触发重启的情况 | 适用场景 |
|---|---|---|---|
no |
从不 | 任何停止 | 一次性任务 |
on-failure |
非零退出、信号终止 | systemctl stop | 大多数Web应用 |
always |
任何停止(包括正常退出) | 无 | 关键服务 |
on-abnormal |
信号、超时、watchdog | 普通非零退出 | 后台工作进程 |
使用 StartLimitInterval 防止重启循环
[Unit]
Description=My Application
StartLimitIntervalSec=60 # Within a 60-second window
StartLimitBurst=3 # Stop trying after 3 restarts
[Service]
Restart=on-failure
RestartSec=10
# If the service restarts more than 3 times in 60 seconds, systemd gives up.
# Run: sudo systemctl reset-failed myapp — before trying to start it again.
测试自动重启
完成服务配置后,请务必在预演环境中模拟崩溃,以确认 systemd 能正确重启进程。检查重启次数和日志输出,确保一切符合预期后再上线生产环境。
# Kill the process — systemd should restart it within RestartSec seconds
sudo kill -9 $(sudo systemctl show myapp -p MainPID | cut -d= -f2)
# Verify it restarted
sudo systemctl status myapp
# Check how many times the service has restarted
sudo systemctl show myapp -p NRestarts