前言:为什么Python比Shell更适合运维自动化
第一章我们掌握了Python基础语法与文件目录操作,能够实现简单的文件类自动化。本章我们正式进入Python运维的核心——专用工具库。
Shell的能力依赖系统命令拼接,而Python通过成熟的第三方库,可以直接获取结构化的系统数据、实现远程批量操作、完成标准化日志处理,代码更简洁、逻辑更清晰、容错性更强。
本章覆盖运维场景最高频的六大类工具库,全部围绕实际工作需求展开,不讲冗余API,只讲生产环境最常用的核心用法,学完就能直接替换掉大量复杂的Shell脚本。
学习收益
- 用subprocess优雅调用系统命令,完全替代os.system()
- 用psutil采集CPU/内存/磁盘指标,无需解析命令输出
- 用requests做接口检测与告警推送,支持完善异常处理
- 用paramiko实现SSH远程执行与SFTP文件分发
- 用logging搭建生产级日志系统,自动轮转与多渠道输出
- 用正则与JSON处理日志和接口数据,比awk/sed更稳定
2.1 命令交互:subprocess模块 调用系统命令与Shell脚本联动
Python无法完全替代系统命令,很多场景下需要调用Shell命令或已有脚本。subprocess是Python官方推荐的进程调用模块,可以执行系统命令、捕获输出与错误码、控制执行状态,是Python与Shell联动的核心桥梁。
2.1.1 核心用法:run()函数
subprocess.run()是最通用的执行方法,功能完整,替代老旧的os.system()、os.popen()。
基础执行
import subprocess
# 执行简单命令,等价于在终端执行 ls -lresult = subprocess.run(["ls", "-l"])规范写法命令与参数分开写在列表中,不开启shell模式,避免注入风险,这是官方推荐的安全写法。
获取命令输出与返回码
运维场景最常用的是捕获命令输出、判断执行是否成功:
import subprocess
# capture_output=True 捕获标准输出和错误输出;text=True 输出转为字符串,默认是字节result = subprocess.run( ["df", "-h", "/"], capture_output=True, text=True)
# 获取返回码,对应Shell的 $?print("返回码:", result.returncode)# 获取标准输出内容print("输出内容:", result.stdout)# 获取错误输出内容print("错误信息:", result.stderr)
# 判断执行是否成功if result.returncode == 0: print("命令执行成功")else: print("命令执行失败")返回码判断最佳实践始终判断
returncode,不能只看输出内容。很多命令即使执行失败也会产生输出,仅凭输出无法准确判断执行结果。建议封装通用的检查函数。
2.1.2 执行Shell语法命令
如果需要用到管道、重定向、通配符等Shell语法,可以开启shell=True,直接传入完整命令字符串:
import subprocess
# 等价于 Shell 的 ps aux | grep nginxresult = subprocess.run( "ps aux | grep nginx", shell=True, capture_output=True, text=True)print(result.stdout)安全风险警告
shell=True存在命令注入风险。如果命令中包含用户输入的变量,禁止使用该模式,必须用列表传参的安全写法。示例攻击代码:# 危险写法 - 永远不要这样做user_input = "test; rm -rf /"subprocess.run(f"echo {user_input}", shell=True) # 灾难性后果
2.1.3 运维典型场景:调用已有Shell脚本
很多团队有存量Shell脚本,不用全部重写,可以用Python调用并接管结果:
import subprocess
# 调用已有的备份脚本,传入参数script_path = "./backup.sh"result = subprocess.run( ["bash", script_path, "/data", "30"], capture_output=True, text=True)
if result.returncode == 0: print("备份脚本执行成功") # 可以进一步解析脚本输出,记录到日志系统else: print("备份脚本执行失败:", result.stderr)2.1.4 高级技巧:超时控制与管道处理
import subprocess
# 超时控制:如果命令超过5秒未完成,自动杀死try: result = subprocess.run( ["sleep", "10"], capture_output=True, text=True, timeout=5 )except subprocess.TimeoutExpired: print("命令执行超时,已强制终止")
# 处理大量输出:使用流式读取,避免内存溢出proc = subprocess.Popen( ["find", "/", "-type", "f"], stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True)
# 逐行处理输出for line in proc.stdout: if "important" in line: print(line.strip())
proc.wait()2.2 系统监控:psutil模块 采集CPU/内存/磁盘/进程全维度指标
Shell里采集系统指标需要拼接ps、df、free、top等命令,再用awk截取字段,繁琐且兼容性差。psutil是Python运维最核心的第三方库,专门用于系统资源与进程管理,直接返回结构化数值,无需解析文本,是监控脚本的首选。
2.2.1 安装与导入
pip3 install psutilimport psutil2.2.2 核心指标采集
1. CPU信息
# 获取CPU逻辑核心数cpu_count = psutil.cpu_count()# 获取CPU物理核心数cpu_physical = psutil.cpu_count(logical=False)
# 获取整体CPU使用率(间隔1秒采样,更准确)cpu_usage = psutil.cpu_percent(interval=1)print(f"CPU使用率:{cpu_usage}%")
# 获取每个核心的使用率cpu_per_core = psutil.cpu_percent(interval=1, percpu=True)for i, usage in enumerate(cpu_per_core): print(f"核心{i}:{usage}%")
# 获取CPU时间分布cpu_times = psutil.cpu_times()print(f"用户态CPU时间:{cpu_times.user}s")print(f"系统态CPU时间:{cpu_times.system}s")2. 内存信息
直接返回所有内存字段,无需像Shell一样用awk截取:
mem = psutil.virtual_memory()
print(f"总内存:{mem.total / 1024**3:.2f} GB")print(f"已用内存:{mem.used / 1024**3:.2f} GB")print(f"可用内存:{mem.available / 1024**3:.2f} GB")print(f"内存使用率:{mem.percent}%")
# Swap交换分区信息swap = psutil.swap_memory()print(f"Swap总大小:{swap.total / 1024**3:.2f} GB")print(f"Swap使用率:{swap.percent}%")3. 磁盘信息
# 获取所有磁盘分区信息partitions = psutil.disk_partitions()for part in partitions: # 获取每个分区的使用情况 usage = psutil.disk_usage(part.mountpoint) print(f"分区 {part.mountpoint}") print(f" 总大小:{usage.total / 1024**3:.2f} GB") print(f" 已用:{usage.used / 1024**3:.2f} GB") print(f" 使用率:{usage.percent}%")
# 获取磁盘IO读写速度(字节/秒)disk_io = psutil.disk_io_counters()print(f"总读取字节数:{disk_io.read_bytes / 1024**3:.2f} GB")print(f"总写入字节数:{disk_io.write_bytes / 1024**3:.2f} GB")4. 进程管理
替代Shell的ps、kill命令,支持按名称筛选、进程信息获取、终止进程:
# 遍历所有进程,获取进程名、PID、内存占用for proc in psutil.process_iter(['pid', 'name', 'memory_percent']): try: print(f"PID:{proc.info['pid']} 名称:{proc.info['name']} 内存占比:{proc.info['memory_percent']:.1f}%") except (psutil.NoSuchProcess, psutil.AccessDenied): pass
# 按名称查找进程nginx_procs = [p for p in psutil.process_iter() if p.name() == "nginx"]print(f"找到{len(nginx_procs)}个nginx进程")
# 获取特定进程的详细信息try: p = psutil.Process(1234) # PID为1234的进程 print(f"进程名:{p.name()}") print(f"执行路径:{p.exe()}") print(f"命令行:{p.cmdline()}") print(f"创建时间:{p.create_time()}") print(f"CPU亲和性:{p.cpu_affinity()}")except psutil.NoSuchProcess: print("进程不存在")
# 终止进程# proc.kill() # 强制杀死,对应 kill -9# proc.terminate() # 优雅终止,对应 kill -152.2.3 实战示例:构建系统健康检查脚本
import psutilimport time
def check_system_health(): """系统健康检查,返回告警列表""" alerts = []
# 检查CPU使用率 cpu_percent = psutil.cpu_percent(interval=1) if cpu_percent > 80: alerts.append(f"⚠️ CPU使用率过高:{cpu_percent}%")
# 检查内存使用率 mem = psutil.virtual_memory() if mem.percent > 85: alerts.append(f"⚠️ 内存使用率过高:{mem.percent}%")
# 检查磁盘使用率 for part in psutil.disk_partitions(): usage = psutil.disk_usage(part.mountpoint) if usage.percent > 90: alerts.append(f"⚠️ 磁盘{part.mountpoint}使用率过高:{usage.percent}%")
# 检查是否有僵尸进程 zombie_count = len([p for p in psutil.process_iter() if p.status() == psutil.STATUS_ZOMBIE]) if zombie_count > 0: alerts.append(f"⚠️ 检测到{zombie_count}个僵尸进程")
return alerts
# 定时检查while True: alerts = check_system_health() if alerts: for alert in alerts: print(alert) else: print("✅ 系统状态正常")
time.sleep(60) # 每分钟检查一次2.2.4 对比Shell的优势
psutil vs Shell命令
功能 Shell方案 psutil方案 获取CPU使用率 top -b -n1 | grep "Cpu"+ awk截取psutil.cpu_percent()获取内存占用 free -h+ awk截取字段psutil.virtual_memory()查询进程信息 ps aux | grep nginx[p for p in psutil.process_iter() if p.name() == "nginx"]数据类型 文本字符串 结构化数值 兼容性 Linux/BSD/Win有差异 跨平台统一 容错性 命令输出格式变化导致脚本失效 API接口稳定
2.3 网络探测:requests模块 HTTP接口检测、数据拉取与告警交互
Shell里做HTTP请求依赖curl命令,解析响应、处理异常、提取JSON数据非常麻烦。requests是Python最流行的HTTP客户端库,语法简洁,功能强大,是接口健康检查、API调用、告警推送的核心工具。
2.3.1 安装与导入
pip3 install requestsimport requests2.3.2 核心基础用法
基础GET请求与响应信息
# 发送GET请求url = "https://www.example.com"response = requests.get(url, timeout=5)
# 获取状态码,对应curl的 %{http_code}print("状态码:", response.status_code)# 获取响应时间(秒)print("响应时间:", response.elapsed.total_seconds())# 获取响应文本内容print("响应内容:", response.text)# 获取响应头print("响应头:", response.headers)
# 判断是否为成功响应if response.ok: print("请求成功")POST请求与参数传递
import requests
# 发送POST请求,传递JSON数据url = "https://api.example.com/submit"data = { "username": "admin", "action": "backup"}
response = requests.post( url, json=data, # 自动转JSON并设置Content-Type timeout=10)
print("响应状态码:", response.status_code)print("响应内容:", response.json()) # 自动解析JSON异常处理
网络请求存在各种异常(超时、连接失败、DNS错误),生产脚本必须捕获异常:
import requests
url = "https://www.example.com"try: response = requests.get(url, timeout=5) response.raise_for_status() # 状态码非2xx时主动抛出异常 print("访问正常,状态码:", response.status_code)except requests.exceptions.Timeout: print("请求超时")except requests.exceptions.ConnectionError: print("连接失败,网络不通")except requests.exceptions.HTTPError as e: print("HTTP错误:", e)except Exception as e: print("未知错误:", e)2.3.3 运维典型场景
场景1:接口健康检查
比Shell的curl写法更清晰,异常处理更完善:
def check_health(url): """检测HTTP接口是否正常,正常返回True,异常返回False""" try: resp = requests.get(url, timeout=3) return resp.status_code == 200 except: return False
if check_health("http://127.0.0.1:8080/health"): print("服务正常")else: print("服务异常,触发告警")场景2:推送告警到企业微信/钉钉
通过webhook地址推送告警消息,是运维自动化告警的常用方式:
def send_wechat_alert(webhook_url, content): """发送企业微信告警""" data = { "msgtype": "text", "text": { "content": content } } try: requests.post(webhook_url, json=data, timeout=5) return True except Exception as e: print(f"发送告警失败:{e}") return False
# 调用示例alert_msg = "【告警】服务器CPU使用率超过90%,请及时处理"send_wechat_alert("你的webhook地址", alert_msg)场景3:API接口数据拉取与解析
def fetch_server_status(api_url, token): """从API拉取服务器状态,返回结构化数据""" headers = { "Authorization": f"Bearer {token}", "Content-Type": "application/json" }
try: response = requests.get(api_url, headers=headers, timeout=10) response.raise_for_status()
# 自动解析JSON data = response.json()
# 提取关键字段 servers = data.get("data", []) for server in servers: print(f"服务器{server['id']}:{server['status']}")
return data except requests.exceptions.HTTPError as e: print(f"API错误:{e.response.status_code}") return None2.3.4 高级技巧:会话管理与连接复用
import requests
# 创建会话对象,复用连接,提高性能session = requests.Session()
# 设置通用headerssession.headers.update({ "User-Agent": "Python-Monitor/1.0", "X-Custom-Header": "value"})
# 多个请求复用连接urls = [ "http://api1.example.com/status", "http://api2.example.com/status", "http://api3.example.com/status"]
for url in urls: try: resp = session.get(url, timeout=5) print(f"{url}: {resp.status_code}") except Exception as e: print(f"{url}: 请求失败 - {e}")
session.close() # 关闭会话2.4 批量运维:paramiko模块 SSH远程执行命令与文件分发
Shell实现批量运维依赖系统ssh、scp命令,需要提前配置密钥,输出解析麻烦。paramiko是Python实现的SSH协议库,可以在代码中直接建立SSH连接、执行命令、传输文件,不依赖系统SSH客户端,是开发批量运维工具的核心。
2.4.1 安装与导入
pip3 install paramikoimport paramiko2.4.2 SSH远程执行命令
def ssh_exec_cmd(host, port, username, password, cmd): """ 远程执行命令,返回 (是否成功, 输出内容, 错误内容) """ ssh = paramiko.SSHClient() # 自动添加主机密钥,避免首次连接提示 ssh.set_missing_host_key_policy(paramiko.AutoAddPolicy())
try: # 建立连接 ssh.connect(hostname=host, port=port, username=username, password=password, timeout=10) # 执行命令 stdin, stdout, stderr = ssh.exec_command(cmd) # 获取输出 out = stdout.read().decode('utf-8') err = stderr.read().decode('utf-8') # 获取返回码 exit_code = stdout.channel.recv_exit_status()
return exit_code == 0, out, err except Exception as e: return False, "", str(e) finally: ssh.close()
# 调用示例success, output, error = ssh_exec_cmd( host="192.168.1.100", port=22, username="root", password="your_password", cmd="df -h /")
if success: print("执行成功:", output)else: print("执行失败:", error)生产环境最佳实践生产环境推荐使用密钥认证替换密码,更安全且无需显式传递密码:
ssh.connect(hostname=host,port=port,username=username,key_filename="/home/user/.ssh/id_rsa", # 私钥文件路径timeout=10)
2.4.3 SFTP文件传输
除了执行命令,还可以实现文件上传下载,替代scp命令:
def sftp_upload(host, port, username, password, local_file, remote_path): """上传本地文件到远程服务器""" transport = paramiko.Transport((host, port)) try: transport.connect(username=username, password=password) sftp = paramiko.SFTPClient.from_transport(transport) sftp.put(local_file, remote_path) sftp.close() print(f"文件{local_file}上传到{host}:{remote_path}成功") return True except Exception as e: print(f"上传失败:{e}") return False finally: transport.close()
def sftp_download(host, port, username, password, remote_file, local_path): """从远程服务器下载文件到本地""" transport = paramiko.Transport((host, port)) try: transport.connect(username=username, password=password) sftp = paramiko.SFTPClient.from_transport(transport) sftp.get(remote_file, local_path) sftp.close() print(f"文件{remote_file}下载到{local_path}成功") return True except Exception as e: print(f"下载失败:{e}") return False finally: transport.close()2.4.4 批量执行思路与进度管理
from concurrent.futures import ThreadPoolExecutor, as_completedimport time
def batch_execute(hosts, cmd, max_workers=5): """ 并发批量执行命令,提高效率 hosts: [{"host": "IP", "port": 22, "user": "root", "pwd": "xxx"}, ...] """ results = {}
def execute_one(server): success, out, err = ssh_exec_cmd( server["host"], server["port"], server["user"], server["pwd"], cmd ) return server["host"], success, out, err
# 使用线程池并发执行 with ThreadPoolExecutor(max_workers=max_workers) as executor: futures = {executor.submit(execute_one, srv): srv for srv in hosts}
for future in as_completed(futures): host, success, out, err = future.result() results[host] = (success, out, err)
# 实时输出进度 if success: print(f"✅ {host} 执行成功") else: print(f"❌ {host} 执行失败:{err[:50]}")
return results
# 使用示例hosts = [ {"host": "192.168.1.101", "port": 22, "user": "root", "pwd": "xxx"}, {"host": "192.168.1.102", "port": 22, "user": "root", "pwd": "xxx"}, {"host": "192.168.1.103", "port": 22, "user": "root", "pwd": "xxx"},]
cmd = "uptime"results = batch_execute(hosts, cmd, max_workers=3)
# 统计结果success_count = sum(1 for _, (ok, _, _) in results.items() if ok)print(f"\n执行完成:{success_count}/{len(hosts)}台服务器成功")相比Shell批量脚本,Python可以更精细地控制每台机器的执行结果、错误处理、超时控制、并发执行,适合开发更专业的批量运维工具。
2.5 日志规范:logging模块 标准化分级日志输出与文件落盘
Shell脚本里我们需要自己封装log_info、log_error函数,功能简陋。Python标准库的logging模块是专业的日志系统,支持日志分级、自定义格式、文件落盘、自动轮转、多渠道输出,是生产脚本的标配。
2.5.1 日志级别
从低到高共5个标准级别,不同级别用于不同场景:
DEBUG:调试信息,开发时使用,生产环境关闭INFO:正常运行信息,比如任务开始、执行成功WARNING:警告信息,不影响运行但需要注意ERROR:错误信息,功能出现异常CRITICAL:严重错误,系统无法继续运行
2.5.2 基础配置与使用
import logging
# 日志基础配置logging.basicConfig( level=logging.INFO, # 设置最低输出级别,低于该级别的日志不显示 format="%(asctime)s [%(levelname)s] %(message)s", # 日志格式 datefmt="%Y-%m-%d %H:%M:%S" # 时间格式)
# 输出不同级别的日志logging.debug("这是调试信息")logging.info("任务开始执行")logging.warning("磁盘使用率超过80%")logging.error("备份执行失败")logging.critical("系统严重故障")运行后会在终端输出带时间、级别的标准化日志,格式统一规范。
2.5.3 同时输出到终端和文件
生产脚本通常需要日志持久化落盘,同时终端也显示:
import loggingfrom logging.handlers import RotatingFileHandler
logger = logging.getLogger()logger.setLevel(logging.INFO)
# 日志格式formatter = logging.Formatter( "%(asctime)s [%(levelname)s] %(name)s - %(message)s", datefmt="%Y-%m-%d %H:%M:%S")
# 1. 终端输出处理器stream_handler = logging.StreamHandler()stream_handler.setFormatter(formatter)logger.addHandler(stream_handler)
# 2. 文件输出处理器(自动轮转,单文件最大10MB,保留5个备份)file_handler = RotatingFileHandler( "run.log", maxBytes=10*1024*1024, backupCount=5, encoding="utf-8")file_handler.setFormatter(formatter)logger.addHandler(file_handler)
# 使用logger.info("服务启动成功")logger.error("数据库连接失败")日志文件轮转机制当单个日志文件达到10MB时,自动重命名为
run.log.1,创建新的run.log。当备份文件超过5个时,自动删除最旧的。这样避免日志文件无限增长。
2.5.4 为不同模块配置独立日志
对于复杂项目,通常按模块配置独立日志记录器:
import logging
def get_logger(name): """获取特定模块的日志记录器""" logger = logging.getLogger(name) if not logger.handlers: handler = logging.FileHandler(f"logs/{name}.log") formatter = logging.Formatter( "%(asctime)s [%(levelname)s] %(message)s", datefmt="%Y-%m-%d %H:%M:%S" ) handler.setFormatter(formatter) logger.addHandler(handler) logger.setLevel(logging.INFO) return logger
# 在不同模块中使用# module_backup.pylogger_backup = get_logger("backup")logger_backup.info("备份开始")
# module_monitor.pylogger_monitor = get_logger("monitor")logger_monitor.error("CPU告警")相比Shell手写的日志函数,logging支持自动日志轮转、避免单文件过大、格式全局统一、支持多渠道输出,是工程化脚本的标准配置。
2.6 数据处理:正则表达式 + JSON解析 日志分析与接口数据提取
2.6.1 正则表达式:re模块
对应Shell的grep、sed正则匹配,Python的re模块功能更强大,适合从日志中提取指定字段。
核心常用方法
import re
log_line = '192.168.1.100 - - [06/Jul/2026:10:30:00 +0800] "GET /index.html HTTP/1.1" 200 1234'
# 1. 匹配判断:是否包含指定模式,对应 grep -qif re.search(r' 200 ', log_line): print("包含200状态码")
# 2. 提取字段:用分组提取IP地址pattern = r'^(\d+\.\d+\.\d+\.\d+)'match = re.match(pattern, log_line)if match: ip = match.group(1) print("提取到IP:", ip)
# 3. 查找所有匹配项,找出日志中所有状态码status_codes = re.findall(r' (\d{3}) ', log_line)print("所有状态码:", status_codes)
# 4. 替换操作,对应 sednew_log = re.sub(r'192\.168', '10\.0', log_line)print("替换后:", new_log)运维典型场景:Nginx日志分析
from collections import Counterimport re
def analyze_nginx_log(log_file): """分析Nginx日志,统计TOP10访问IP、状态码分布、热点URL"""
ip_list = [] status_dict = Counter() url_dict = Counter()
with open(log_file, "r", encoding="utf-8") as f: for line in f: # 提取IP:正则匹配日志行的IP字段 ip_match = re.match(r'^(\d+\.\d+\.\d+\.\d+)', line) if ip_match: ip_list.append(ip_match.group(1))
# 提取状态码 status_match = re.search(r' (\d{3}) ', line) if status_match: status_dict[status_match.group(1)] += 1
# 提取URL url_match = re.search(r'"[A-Z]+ ([^\s]+) HTTP', line) if url_match: url_dict[url_match.group(1)] += 1
print("=" * 50) print("【TOP 10 访问IP】") for ip, count in Counter(ip_list).most_common(10): print(f" {ip:20} {count:6}次")
print("\n【状态码分布】") for status, count in sorted(status_dict.items()): print(f" HTTP {status}: {count}次")
print("\n【TOP 10 热点URL】") for url, count in url_dict.most_common(10): print(f" {url:40} {count:6}次")
# 使用analyze_nginx_log("/var/log/nginx/access.log")2.6.2 JSON数据解析
Shell处理JSON需要额外安装jq工具,而Python原生支持JSON解析,处理接口返回数据、配置文件非常方便。
import json
# JSON字符串转Python字典json_str = '{"code": 200, "data": {"status": "running", "cpu": 45}}'data = json.loads(json_str)
# 直接按键取值,无需正则截取print("状态码:", data["code"])print("服务状态:", data["data"]["status"])print("CPU使用率:", data["data"]["cpu"])
# Python字典转JSON字符串result = {"success": True, "message": "执行完成"}json_result = json.dumps(result, ensure_ascii=False, indent=2)print(json_result)
# 安全处理不存在的键data = {"code": 200}cpu = data.get("cpu", "未知") # 键不存在时返回默认值print("CPU:", cpu)2.6.3 实战场景:读取配置文件并校验
import jsonimport logging
logger = logging.getLogger(__name__)
def load_config(config_file): """加载JSON配置文件,支持容错""" try: with open(config_file, "r", encoding="utf-8") as f: config = json.load(f)
# 配置校验 required_keys = ["servers", "timeout", "retry"] for key in required_keys: if key not in config: raise ValueError(f"配置文件缺少必须字段:{key}")
logger.info(f"配置文件加载成功,包含{len(config['servers'])}个服务器") return config except json.JSONDecodeError as e: logger.error(f"配置文件格式错误:{e}") return None except Exception as e: logger.error(f"加载配置文件失败:{e}") return None
# 配置文件 config.json"""{ "servers": [ {"host": "192.168.1.100", "port": 22}, {"host": "192.168.1.101", "port": 22} ], "timeout": 30, "retry": 3}"""
config = load_config("config.json")if config: for server in config["servers"]: print(f"连接到{server['host']}:{server['port']}")这是Python相比Shell的巨大优势之一:处理结构化数据(JSON、YAML、XML)天然支持,不用依赖外部工具,代码更稳定、容错性更强。
2.7 本章小结 + 综合实战
2.7.1 本章知识点速记
六大工具库核心用法
模块 核心功能 最常用API 适用场景 subprocess 命令执行 run(...capture_output=True)调用系统命令、调用Shell脚本 psutil 系统监控 cpu_percent(),virtual_memory(),disk_partitions()CPU/内存/磁盘监控、进程管理 requests HTTP客户端 get(),post(),raise_for_status()接口检测、数据拉取、告警推送 paramiko SSH协议 SSHClient.connect(),exec_command(),SFTP远程执行、文件分发、批量运维 logging 日志系统 basicConfig(),RotatingFileHandler生产脚本日志、多级别输出 re + json 数据处理 re.findall(),json.loads()日志解析、配置文件、接口数据提取
2.7.2 综合实战:监控告警系统
集成本章所有工具库,构建一个完整的服务器监控告警脚本:
import psutilimport requestsimport subprocessimport loggingfrom logging.handlers import RotatingFileHandlerimport timefrom datetime import datetime
# ============= 日志配置 =============logger = logging.getLogger()logger.setLevel(logging.INFO)
formatter = logging.Formatter( "%(asctime)s [%(levelname)s] %(message)s", datefmt="%Y-%m-%d %H:%M:%S")
stream_handler = logging.StreamHandler()stream_handler.setFormatter(formatter)logger.addHandler(stream_handler)
file_handler = RotatingFileHandler( "monitor.log", maxBytes=10*1024*1024, backupCount=5, encoding="utf-8")file_handler.setFormatter(formatter)logger.addHandler(file_handler)
# ============= 告警推送 =============def send_alert(title, content, webhook_url): """推送告警到企业微信""" try: data = { "msgtype": "markdown", "markdown": { "content": f"**{title}**\n\n{content}" } } requests.post(webhook_url, json=data, timeout=5) logger.info(f"告警已推送:{title}") except Exception as e: logger.error(f"告警推送失败:{e}")
# ============= 系统检查 =============def check_system_health(thresholds, webhook_url): """综合系统检查,超过阈值触发告警""" alerts = []
# 检查CPU cpu_percent = psutil.cpu_percent(interval=1) if cpu_percent > thresholds["cpu"]: msg = f"CPU使用率告警:{cpu_percent}% (阈值:{thresholds['cpu']}%)" alerts.append(("CPU告警", msg)) logger.warning(msg)
# 检查内存 mem = psutil.virtual_memory() if mem.percent > thresholds["memory"]: msg = f"内存使用率告警:{mem.percent}% (阈值:{thresholds['memory']}%)" alerts.append(("内存告警", msg)) logger.warning(msg)
# 检查磁盘 for part in psutil.disk_partitions(): usage = psutil.disk_usage(part.mountpoint) if usage.percent > thresholds["disk"]: msg = f"磁盘告警 {part.mountpoint}:{usage.percent}% (阈值:{thresholds['disk']}%)" alerts.append(("磁盘告警", msg)) logger.warning(msg)
# 检查进程 try: # 检查关键服务是否运行 result = subprocess.run( ["pgrep", "-a", "nginx"], capture_output=True, text=True ) if result.returncode != 0: msg = "Nginx服务未运行" alerts.append(("服务告警", msg)) logger.error(msg) except Exception as e: logger.error(f"进程检查异常:{e}")
# 推送告警 for title, msg in alerts: send_alert(title, msg, webhook_url)
return len(alerts) > 0
# ============= 定时监控 =============def monitor_loop(thresholds, webhook_url, interval=60): """启动监控循环""" logger.info("监控系统启动")
check_count = 0 alert_count = 0
try: while True: check_count += 1
if check_system_health(thresholds, webhook_url): alert_count += 1
# 每小时打印一次统计 if check_count % 60 == 0: logger.info(f"检查统计:总检查{check_count}次,告警{alert_count}次")
time.sleep(interval) except KeyboardInterrupt: logger.info("监控系统已停止") except Exception as e: logger.critical(f"监控系统异常退出:{e}")
# ============= 启动程序 =============if __name__ == "__main__": # 监控阈值配置 thresholds = { "cpu": 80, # CPU超过80%告警 "memory": 85, # 内存超过85%告警 "disk": 90 # 磁盘超过90%告警 }
# 企业微信webhook地址(需要替换为实际地址) webhook_url = "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=xxx"
# 启动监控 monitor_loop(thresholds, webhook_url, interval=60)这个脚本展示了如何整合:
psutil采集系统指标subprocess执行系统命令检查进程requests推送告警logging记录所有操作和告警- 完整的异常处理和日志记录