5483 字
27 分钟
Python第二章 运维核心工具库:高频场景专用模块

前言:为什么Python比Shell更适合运维自动化#

第一章我们掌握了Python基础语法与文件目录操作,能够实现简单的文件类自动化。本章我们正式进入Python运维的核心——专用工具库。

Shell的能力依赖系统命令拼接,而Python通过成熟的第三方库,可以直接获取结构化的系统数据、实现远程批量操作、完成标准化日志处理,代码更简洁、逻辑更清晰、容错性更强。

本章覆盖运维场景最高频的六大类工具库,全部围绕实际工作需求展开,不讲冗余API,只讲生产环境最常用的核心用法,学完就能直接替换掉大量复杂的Shell脚本。

学习收益
  • 用subprocess优雅调用系统命令,完全替代os.system()
  • 用psutil采集CPU/内存/磁盘指标,无需解析命令输出
  • 用requests做接口检测与告警推送,支持完善异常处理
  • 用paramiko实现SSH远程执行与SFTP文件分发
  • 用logging搭建生产级日志系统,自动轮转与多渠道输出
  • 用正则与JSON处理日志和接口数据,比awk/sed更稳定

2.1 命令交互:subprocess模块 调用系统命令与Shell脚本联动#

Python无法完全替代系统命令,很多场景下需要调用Shell命令或已有脚本。subprocess是Python官方推荐的进程调用模块,可以执行系统命令、捕获输出与错误码、控制执行状态,是Python与Shell联动的核心桥梁。

2.1.1 核心用法:run()函数#

subprocess.run()是最通用的执行方法,功能完整,替代老旧的os.system()、os.popen()。

基础执行#

import subprocess
# 执行简单命令,等价于在终端执行 ls -l
result = subprocess.run(["ls", "-l"])
规范写法

命令与参数分开写在列表中,不开启shell模式,避免注入风险,这是官方推荐的安全写法。

获取命令输出与返回码#

运维场景最常用的是捕获命令输出、判断执行是否成功:

import subprocess
# capture_output=True 捕获标准输出和错误输出;text=True 输出转为字符串,默认是字节
result = subprocess.run(
["df", "-h", "/"],
capture_output=True,
text=True
)
# 获取返回码,对应Shell的 $?
print("返回码:", result.returncode)
# 获取标准输出内容
print("输出内容:", result.stdout)
# 获取错误输出内容
print("错误信息:", result.stderr)
# 判断执行是否成功
if result.returncode == 0:
print("命令执行成功")
else:
print("命令执行失败")
返回码判断最佳实践

始终判断returncode,不能只看输出内容。很多命令即使执行失败也会产生输出,仅凭输出无法准确判断执行结果。建议封装通用的检查函数。

2.1.2 执行Shell语法命令#

如果需要用到管道、重定向、通配符等Shell语法,可以开启shell=True,直接传入完整命令字符串:

import subprocess
# 等价于 Shell 的 ps aux | grep nginx
result = subprocess.run(
"ps aux | grep nginx",
shell=True,
capture_output=True,
text=True
)
print(result.stdout)
安全风险警告

shell=True存在命令注入风险。如果命令中包含用户输入的变量,禁止使用该模式,必须用列表传参的安全写法。示例攻击代码:

# 危险写法 - 永远不要这样做
user_input = "test; rm -rf /"
subprocess.run(f"echo {user_input}", shell=True) # 灾难性后果

2.1.3 运维典型场景:调用已有Shell脚本#

很多团队有存量Shell脚本,不用全部重写,可以用Python调用并接管结果:

import subprocess
# 调用已有的备份脚本,传入参数
script_path = "./backup.sh"
result = subprocess.run(
["bash", script_path, "/data", "30"],
capture_output=True,
text=True
)
if result.returncode == 0:
print("备份脚本执行成功")
# 可以进一步解析脚本输出,记录到日志系统
else:
print("备份脚本执行失败:", result.stderr)

2.1.4 高级技巧:超时控制与管道处理#

import subprocess
# 超时控制:如果命令超过5秒未完成,自动杀死
try:
result = subprocess.run(
["sleep", "10"],
capture_output=True,
text=True,
timeout=5
)
except subprocess.TimeoutExpired:
print("命令执行超时,已强制终止")
# 处理大量输出:使用流式读取,避免内存溢出
proc = subprocess.Popen(
["find", "/", "-type", "f"],
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
text=True
)
# 逐行处理输出
for line in proc.stdout:
if "important" in line:
print(line.strip())
proc.wait()

2.2 系统监控:psutil模块 采集CPU/内存/磁盘/进程全维度指标#

Shell里采集系统指标需要拼接ps、df、free、top等命令,再用awk截取字段,繁琐且兼容性差。psutil是Python运维最核心的第三方库,专门用于系统资源与进程管理,直接返回结构化数值,无需解析文本,是监控脚本的首选。

2.2.1 安装与导入#

Terminal window
pip3 install psutil
import psutil

2.2.2 核心指标采集#

1. CPU信息#

# 获取CPU逻辑核心数
cpu_count = psutil.cpu_count()
# 获取CPU物理核心数
cpu_physical = psutil.cpu_count(logical=False)
# 获取整体CPU使用率(间隔1秒采样,更准确)
cpu_usage = psutil.cpu_percent(interval=1)
print(f"CPU使用率:{cpu_usage}%")
# 获取每个核心的使用率
cpu_per_core = psutil.cpu_percent(interval=1, percpu=True)
for i, usage in enumerate(cpu_per_core):
print(f"核心{i}:{usage}%")
# 获取CPU时间分布
cpu_times = psutil.cpu_times()
print(f"用户态CPU时间:{cpu_times.user}s")
print(f"系统态CPU时间:{cpu_times.system}s")

2. 内存信息#

直接返回所有内存字段,无需像Shell一样用awk截取:

mem = psutil.virtual_memory()
print(f"总内存:{mem.total / 1024**3:.2f} GB")
print(f"已用内存:{mem.used / 1024**3:.2f} GB")
print(f"可用内存:{mem.available / 1024**3:.2f} GB")
print(f"内存使用率:{mem.percent}%")
# Swap交换分区信息
swap = psutil.swap_memory()
print(f"Swap总大小:{swap.total / 1024**3:.2f} GB")
print(f"Swap使用率:{swap.percent}%")

3. 磁盘信息#

# 获取所有磁盘分区信息
partitions = psutil.disk_partitions()
for part in partitions:
# 获取每个分区的使用情况
usage = psutil.disk_usage(part.mountpoint)
print(f"分区 {part.mountpoint}")
print(f" 总大小:{usage.total / 1024**3:.2f} GB")
print(f" 已用:{usage.used / 1024**3:.2f} GB")
print(f" 使用率:{usage.percent}%")
# 获取磁盘IO读写速度(字节/秒)
disk_io = psutil.disk_io_counters()
print(f"总读取字节数:{disk_io.read_bytes / 1024**3:.2f} GB")
print(f"总写入字节数:{disk_io.write_bytes / 1024**3:.2f} GB")

4. 进程管理#

替代Shell的ps、kill命令,支持按名称筛选、进程信息获取、终止进程:

# 遍历所有进程,获取进程名、PID、内存占用
for proc in psutil.process_iter(['pid', 'name', 'memory_percent']):
try:
print(f"PID:{proc.info['pid']} 名称:{proc.info['name']} 内存占比:{proc.info['memory_percent']:.1f}%")
except (psutil.NoSuchProcess, psutil.AccessDenied):
pass
# 按名称查找进程
nginx_procs = [p for p in psutil.process_iter() if p.name() == "nginx"]
print(f"找到{len(nginx_procs)}个nginx进程")
# 获取特定进程的详细信息
try:
p = psutil.Process(1234) # PID为1234的进程
print(f"进程名:{p.name()}")
print(f"执行路径:{p.exe()}")
print(f"命令行:{p.cmdline()}")
print(f"创建时间:{p.create_time()}")
print(f"CPU亲和性:{p.cpu_affinity()}")
except psutil.NoSuchProcess:
print("进程不存在")
# 终止进程
# proc.kill() # 强制杀死,对应 kill -9
# proc.terminate() # 优雅终止,对应 kill -15

2.2.3 实战示例:构建系统健康检查脚本#

import psutil
import time
def check_system_health():
"""系统健康检查,返回告警列表"""
alerts = []
# 检查CPU使用率
cpu_percent = psutil.cpu_percent(interval=1)
if cpu_percent > 80:
alerts.append(f"⚠️ CPU使用率过高:{cpu_percent}%")
# 检查内存使用率
mem = psutil.virtual_memory()
if mem.percent > 85:
alerts.append(f"⚠️ 内存使用率过高:{mem.percent}%")
# 检查磁盘使用率
for part in psutil.disk_partitions():
usage = psutil.disk_usage(part.mountpoint)
if usage.percent > 90:
alerts.append(f"⚠️ 磁盘{part.mountpoint}使用率过高:{usage.percent}%")
# 检查是否有僵尸进程
zombie_count = len([p for p in psutil.process_iter() if p.status() == psutil.STATUS_ZOMBIE])
if zombie_count > 0:
alerts.append(f"⚠️ 检测到{zombie_count}个僵尸进程")
return alerts
# 定时检查
while True:
alerts = check_system_health()
if alerts:
for alert in alerts:
print(alert)
else:
print("✅ 系统状态正常")
time.sleep(60) # 每分钟检查一次

2.2.4 对比Shell的优势#

psutil vs Shell命令
功能Shell方案psutil方案
获取CPU使用率top -b -n1 | grep "Cpu" + awk截取psutil.cpu_percent()
获取内存占用free -h + awk截取字段psutil.virtual_memory()
查询进程信息ps aux | grep nginx[p for p in psutil.process_iter() if p.name() == "nginx"]
数据类型文本字符串结构化数值
兼容性Linux/BSD/Win有差异跨平台统一
容错性命令输出格式变化导致脚本失效API接口稳定

2.3 网络探测:requests模块 HTTP接口检测、数据拉取与告警交互#

Shell里做HTTP请求依赖curl命令,解析响应、处理异常、提取JSON数据非常麻烦。requests是Python最流行的HTTP客户端库,语法简洁,功能强大,是接口健康检查、API调用、告警推送的核心工具。

2.3.1 安装与导入#

Terminal window
pip3 install requests
import requests

2.3.2 核心基础用法#

基础GET请求与响应信息#

# 发送GET请求
url = "https://www.example.com"
response = requests.get(url, timeout=5)
# 获取状态码,对应curl的 %{http_code}
print("状态码:", response.status_code)
# 获取响应时间(秒)
print("响应时间:", response.elapsed.total_seconds())
# 获取响应文本内容
print("响应内容:", response.text)
# 获取响应头
print("响应头:", response.headers)
# 判断是否为成功响应
if response.ok:
print("请求成功")

POST请求与参数传递#

import requests
# 发送POST请求,传递JSON数据
url = "https://api.example.com/submit"
data = {
"username": "admin",
"action": "backup"
}
response = requests.post(
url,
json=data, # 自动转JSON并设置Content-Type
timeout=10
)
print("响应状态码:", response.status_code)
print("响应内容:", response.json()) # 自动解析JSON

异常处理#

网络请求存在各种异常(超时、连接失败、DNS错误),生产脚本必须捕获异常:

import requests
url = "https://www.example.com"
try:
response = requests.get(url, timeout=5)
response.raise_for_status() # 状态码非2xx时主动抛出异常
print("访问正常,状态码:", response.status_code)
except requests.exceptions.Timeout:
print("请求超时")
except requests.exceptions.ConnectionError:
print("连接失败,网络不通")
except requests.exceptions.HTTPError as e:
print("HTTP错误:", e)
except Exception as e:
print("未知错误:", e)

2.3.3 运维典型场景#

场景1:接口健康检查#

比Shell的curl写法更清晰,异常处理更完善:

def check_health(url):
"""检测HTTP接口是否正常,正常返回True,异常返回False"""
try:
resp = requests.get(url, timeout=3)
return resp.status_code == 200
except:
return False
if check_health("http://127.0.0.1:8080/health"):
print("服务正常")
else:
print("服务异常,触发告警")

场景2:推送告警到企业微信/钉钉#

通过webhook地址推送告警消息,是运维自动化告警的常用方式:

def send_wechat_alert(webhook_url, content):
"""发送企业微信告警"""
data = {
"msgtype": "text",
"text": {
"content": content
}
}
try:
requests.post(webhook_url, json=data, timeout=5)
return True
except Exception as e:
print(f"发送告警失败:{e}")
return False
# 调用示例
alert_msg = "【告警】服务器CPU使用率超过90%,请及时处理"
send_wechat_alert("你的webhook地址", alert_msg)

场景3:API接口数据拉取与解析#

def fetch_server_status(api_url, token):
"""从API拉取服务器状态,返回结构化数据"""
headers = {
"Authorization": f"Bearer {token}",
"Content-Type": "application/json"
}
try:
response = requests.get(api_url, headers=headers, timeout=10)
response.raise_for_status()
# 自动解析JSON
data = response.json()
# 提取关键字段
servers = data.get("data", [])
for server in servers:
print(f"服务器{server['id']}:{server['status']}")
return data
except requests.exceptions.HTTPError as e:
print(f"API错误:{e.response.status_code}")
return None

2.3.4 高级技巧:会话管理与连接复用#

import requests
# 创建会话对象,复用连接,提高性能
session = requests.Session()
# 设置通用headers
session.headers.update({
"User-Agent": "Python-Monitor/1.0",
"X-Custom-Header": "value"
})
# 多个请求复用连接
urls = [
"http://api1.example.com/status",
"http://api2.example.com/status",
"http://api3.example.com/status"
]
for url in urls:
try:
resp = session.get(url, timeout=5)
print(f"{url}: {resp.status_code}")
except Exception as e:
print(f"{url}: 请求失败 - {e}")
session.close() # 关闭会话

2.4 批量运维:paramiko模块 SSH远程执行命令与文件分发#

Shell实现批量运维依赖系统ssh、scp命令,需要提前配置密钥,输出解析麻烦。paramiko是Python实现的SSH协议库,可以在代码中直接建立SSH连接、执行命令、传输文件,不依赖系统SSH客户端,是开发批量运维工具的核心。

2.4.1 安装与导入#

Terminal window
pip3 install paramiko
import paramiko

2.4.2 SSH远程执行命令#

def ssh_exec_cmd(host, port, username, password, cmd):
"""
远程执行命令,返回 (是否成功, 输出内容, 错误内容)
"""
ssh = paramiko.SSHClient()
# 自动添加主机密钥,避免首次连接提示
ssh.set_missing_host_key_policy(paramiko.AutoAddPolicy())
try:
# 建立连接
ssh.connect(hostname=host, port=port, username=username, password=password, timeout=10)
# 执行命令
stdin, stdout, stderr = ssh.exec_command(cmd)
# 获取输出
out = stdout.read().decode('utf-8')
err = stderr.read().decode('utf-8')
# 获取返回码
exit_code = stdout.channel.recv_exit_status()
return exit_code == 0, out, err
except Exception as e:
return False, "", str(e)
finally:
ssh.close()
# 调用示例
success, output, error = ssh_exec_cmd(
host="192.168.1.100",
port=22,
username="root",
password="your_password",
cmd="df -h /"
)
if success:
print("执行成功:", output)
else:
print("执行失败:", error)
生产环境最佳实践

生产环境推荐使用密钥认证替换密码,更安全且无需显式传递密码:

ssh.connect(
hostname=host,
port=port,
username=username,
key_filename="/home/user/.ssh/id_rsa", # 私钥文件路径
timeout=10
)

2.4.3 SFTP文件传输#

除了执行命令,还可以实现文件上传下载,替代scp命令:

def sftp_upload(host, port, username, password, local_file, remote_path):
"""上传本地文件到远程服务器"""
transport = paramiko.Transport((host, port))
try:
transport.connect(username=username, password=password)
sftp = paramiko.SFTPClient.from_transport(transport)
sftp.put(local_file, remote_path)
sftp.close()
print(f"文件{local_file}上传到{host}:{remote_path}成功")
return True
except Exception as e:
print(f"上传失败:{e}")
return False
finally:
transport.close()
def sftp_download(host, port, username, password, remote_file, local_path):
"""从远程服务器下载文件到本地"""
transport = paramiko.Transport((host, port))
try:
transport.connect(username=username, password=password)
sftp = paramiko.SFTPClient.from_transport(transport)
sftp.get(remote_file, local_path)
sftp.close()
print(f"文件{remote_file}下载到{local_path}成功")
return True
except Exception as e:
print(f"下载失败:{e}")
return False
finally:
transport.close()

2.4.4 批量执行思路与进度管理#

from concurrent.futures import ThreadPoolExecutor, as_completed
import time
def batch_execute(hosts, cmd, max_workers=5):
"""
并发批量执行命令,提高效率
hosts: [{"host": "IP", "port": 22, "user": "root", "pwd": "xxx"}, ...]
"""
results = {}
def execute_one(server):
success, out, err = ssh_exec_cmd(
server["host"],
server["port"],
server["user"],
server["pwd"],
cmd
)
return server["host"], success, out, err
# 使用线程池并发执行
with ThreadPoolExecutor(max_workers=max_workers) as executor:
futures = {executor.submit(execute_one, srv): srv for srv in hosts}
for future in as_completed(futures):
host, success, out, err = future.result()
results[host] = (success, out, err)
# 实时输出进度
if success:
print(f"✅ {host} 执行成功")
else:
print(f"❌ {host} 执行失败:{err[:50]}")
return results
# 使用示例
hosts = [
{"host": "192.168.1.101", "port": 22, "user": "root", "pwd": "xxx"},
{"host": "192.168.1.102", "port": 22, "user": "root", "pwd": "xxx"},
{"host": "192.168.1.103", "port": 22, "user": "root", "pwd": "xxx"},
]
cmd = "uptime"
results = batch_execute(hosts, cmd, max_workers=3)
# 统计结果
success_count = sum(1 for _, (ok, _, _) in results.items() if ok)
print(f"\n执行完成:{success_count}/{len(hosts)}台服务器成功")

相比Shell批量脚本,Python可以更精细地控制每台机器的执行结果、错误处理、超时控制、并发执行,适合开发更专业的批量运维工具。


2.5 日志规范:logging模块 标准化分级日志输出与文件落盘#

Shell脚本里我们需要自己封装log_info、log_error函数,功能简陋。Python标准库的logging模块是专业的日志系统,支持日志分级、自定义格式、文件落盘、自动轮转、多渠道输出,是生产脚本的标配。

2.5.1 日志级别#

从低到高共5个标准级别,不同级别用于不同场景:

  • DEBUG:调试信息,开发时使用,生产环境关闭
  • INFO:正常运行信息,比如任务开始、执行成功
  • WARNING:警告信息,不影响运行但需要注意
  • ERROR:错误信息,功能出现异常
  • CRITICAL:严重错误,系统无法继续运行

2.5.2 基础配置与使用#

import logging
# 日志基础配置
logging.basicConfig(
level=logging.INFO, # 设置最低输出级别,低于该级别的日志不显示
format="%(asctime)s [%(levelname)s] %(message)s", # 日志格式
datefmt="%Y-%m-%d %H:%M:%S" # 时间格式
)
# 输出不同级别的日志
logging.debug("这是调试信息")
logging.info("任务开始执行")
logging.warning("磁盘使用率超过80%")
logging.error("备份执行失败")
logging.critical("系统严重故障")

运行后会在终端输出带时间、级别的标准化日志,格式统一规范。

2.5.3 同时输出到终端和文件#

生产脚本通常需要日志持久化落盘,同时终端也显示:

import logging
from logging.handlers import RotatingFileHandler
logger = logging.getLogger()
logger.setLevel(logging.INFO)
# 日志格式
formatter = logging.Formatter(
"%(asctime)s [%(levelname)s] %(name)s - %(message)s",
datefmt="%Y-%m-%d %H:%M:%S"
)
# 1. 终端输出处理器
stream_handler = logging.StreamHandler()
stream_handler.setFormatter(formatter)
logger.addHandler(stream_handler)
# 2. 文件输出处理器(自动轮转,单文件最大10MB,保留5个备份)
file_handler = RotatingFileHandler(
"run.log",
maxBytes=10*1024*1024,
backupCount=5,
encoding="utf-8"
)
file_handler.setFormatter(formatter)
logger.addHandler(file_handler)
# 使用
logger.info("服务启动成功")
logger.error("数据库连接失败")
日志文件轮转机制

当单个日志文件达到10MB时,自动重命名为run.log.1,创建新的run.log。当备份文件超过5个时,自动删除最旧的。这样避免日志文件无限增长。

2.5.4 为不同模块配置独立日志#

对于复杂项目,通常按模块配置独立日志记录器:

import logging
def get_logger(name):
"""获取特定模块的日志记录器"""
logger = logging.getLogger(name)
if not logger.handlers:
handler = logging.FileHandler(f"logs/{name}.log")
formatter = logging.Formatter(
"%(asctime)s [%(levelname)s] %(message)s",
datefmt="%Y-%m-%d %H:%M:%S"
)
handler.setFormatter(formatter)
logger.addHandler(handler)
logger.setLevel(logging.INFO)
return logger
# 在不同模块中使用
# module_backup.py
logger_backup = get_logger("backup")
logger_backup.info("备份开始")
# module_monitor.py
logger_monitor = get_logger("monitor")
logger_monitor.error("CPU告警")

相比Shell手写的日志函数,logging支持自动日志轮转、避免单文件过大、格式全局统一、支持多渠道输出,是工程化脚本的标准配置。


2.6 数据处理:正则表达式 + JSON解析 日志分析与接口数据提取#

2.6.1 正则表达式:re模块#

对应Shell的grep、sed正则匹配,Python的re模块功能更强大,适合从日志中提取指定字段。

核心常用方法#

import re
log_line = '192.168.1.100 - - [06/Jul/2026:10:30:00 +0800] "GET /index.html HTTP/1.1" 200 1234'
# 1. 匹配判断:是否包含指定模式,对应 grep -q
if re.search(r' 200 ', log_line):
print("包含200状态码")
# 2. 提取字段:用分组提取IP地址
pattern = r'^(\d+\.\d+\.\d+\.\d+)'
match = re.match(pattern, log_line)
if match:
ip = match.group(1)
print("提取到IP:", ip)
# 3. 查找所有匹配项,找出日志中所有状态码
status_codes = re.findall(r' (\d{3}) ', log_line)
print("所有状态码:", status_codes)
# 4. 替换操作,对应 sed
new_log = re.sub(r'192\.168', '10\.0', log_line)
print("替换后:", new_log)

运维典型场景:Nginx日志分析#

from collections import Counter
import re
def analyze_nginx_log(log_file):
"""分析Nginx日志,统计TOP10访问IP、状态码分布、热点URL"""
ip_list = []
status_dict = Counter()
url_dict = Counter()
with open(log_file, "r", encoding="utf-8") as f:
for line in f:
# 提取IP:正则匹配日志行的IP字段
ip_match = re.match(r'^(\d+\.\d+\.\d+\.\d+)', line)
if ip_match:
ip_list.append(ip_match.group(1))
# 提取状态码
status_match = re.search(r' (\d{3}) ', line)
if status_match:
status_dict[status_match.group(1)] += 1
# 提取URL
url_match = re.search(r'"[A-Z]+ ([^\s]+) HTTP', line)
if url_match:
url_dict[url_match.group(1)] += 1
print("=" * 50)
print("【TOP 10 访问IP】")
for ip, count in Counter(ip_list).most_common(10):
print(f" {ip:20} {count:6}次")
print("\n【状态码分布】")
for status, count in sorted(status_dict.items()):
print(f" HTTP {status}: {count}次")
print("\n【TOP 10 热点URL】")
for url, count in url_dict.most_common(10):
print(f" {url:40} {count:6}次")
# 使用
analyze_nginx_log("/var/log/nginx/access.log")

2.6.2 JSON数据解析#

Shell处理JSON需要额外安装jq工具,而Python原生支持JSON解析,处理接口返回数据、配置文件非常方便。

import json
# JSON字符串转Python字典
json_str = '{"code": 200, "data": {"status": "running", "cpu": 45}}'
data = json.loads(json_str)
# 直接按键取值,无需正则截取
print("状态码:", data["code"])
print("服务状态:", data["data"]["status"])
print("CPU使用率:", data["data"]["cpu"])
# Python字典转JSON字符串
result = {"success": True, "message": "执行完成"}
json_result = json.dumps(result, ensure_ascii=False, indent=2)
print(json_result)
# 安全处理不存在的键
data = {"code": 200}
cpu = data.get("cpu", "未知") # 键不存在时返回默认值
print("CPU:", cpu)

2.6.3 实战场景:读取配置文件并校验#

import json
import logging
logger = logging.getLogger(__name__)
def load_config(config_file):
"""加载JSON配置文件,支持容错"""
try:
with open(config_file, "r", encoding="utf-8") as f:
config = json.load(f)
# 配置校验
required_keys = ["servers", "timeout", "retry"]
for key in required_keys:
if key not in config:
raise ValueError(f"配置文件缺少必须字段:{key}")
logger.info(f"配置文件加载成功,包含{len(config['servers'])}个服务器")
return config
except json.JSONDecodeError as e:
logger.error(f"配置文件格式错误:{e}")
return None
except Exception as e:
logger.error(f"加载配置文件失败:{e}")
return None
# 配置文件 config.json
"""
{
"servers": [
{"host": "192.168.1.100", "port": 22},
{"host": "192.168.1.101", "port": 22}
],
"timeout": 30,
"retry": 3
}
"""
config = load_config("config.json")
if config:
for server in config["servers"]:
print(f"连接到{server['host']}:{server['port']}")

这是Python相比Shell的巨大优势之一:处理结构化数据(JSON、YAML、XML)天然支持,不用依赖外部工具,代码更稳定、容错性更强。


2.7 本章小结 + 综合实战#

2.7.1 本章知识点速记#

六大工具库核心用法
模块核心功能最常用API适用场景
subprocess命令执行run(...capture_output=True)调用系统命令、调用Shell脚本
psutil系统监控cpu_percent(), virtual_memory(), disk_partitions()CPU/内存/磁盘监控、进程管理
requestsHTTP客户端get(), post(), raise_for_status()接口检测、数据拉取、告警推送
paramikoSSH协议SSHClient.connect(), exec_command(), SFTP远程执行、文件分发、批量运维
logging日志系统basicConfig(), RotatingFileHandler生产脚本日志、多级别输出
re + json数据处理re.findall(), json.loads()日志解析、配置文件、接口数据提取

2.7.2 综合实战:监控告警系统#

集成本章所有工具库,构建一个完整的服务器监控告警脚本:

import psutil
import requests
import subprocess
import logging
from logging.handlers import RotatingFileHandler
import time
from datetime import datetime
# ============= 日志配置 =============
logger = logging.getLogger()
logger.setLevel(logging.INFO)
formatter = logging.Formatter(
"%(asctime)s [%(levelname)s] %(message)s",
datefmt="%Y-%m-%d %H:%M:%S"
)
stream_handler = logging.StreamHandler()
stream_handler.setFormatter(formatter)
logger.addHandler(stream_handler)
file_handler = RotatingFileHandler(
"monitor.log",
maxBytes=10*1024*1024,
backupCount=5,
encoding="utf-8"
)
file_handler.setFormatter(formatter)
logger.addHandler(file_handler)
# ============= 告警推送 =============
def send_alert(title, content, webhook_url):
"""推送告警到企业微信"""
try:
data = {
"msgtype": "markdown",
"markdown": {
"content": f"**{title}**\n\n{content}"
}
}
requests.post(webhook_url, json=data, timeout=5)
logger.info(f"告警已推送:{title}")
except Exception as e:
logger.error(f"告警推送失败:{e}")
# ============= 系统检查 =============
def check_system_health(thresholds, webhook_url):
"""综合系统检查,超过阈值触发告警"""
alerts = []
# 检查CPU
cpu_percent = psutil.cpu_percent(interval=1)
if cpu_percent > thresholds["cpu"]:
msg = f"CPU使用率告警:{cpu_percent}% (阈值:{thresholds['cpu']}%)"
alerts.append(("CPU告警", msg))
logger.warning(msg)
# 检查内存
mem = psutil.virtual_memory()
if mem.percent > thresholds["memory"]:
msg = f"内存使用率告警:{mem.percent}% (阈值:{thresholds['memory']}%)"
alerts.append(("内存告警", msg))
logger.warning(msg)
# 检查磁盘
for part in psutil.disk_partitions():
usage = psutil.disk_usage(part.mountpoint)
if usage.percent > thresholds["disk"]:
msg = f"磁盘告警 {part.mountpoint}:{usage.percent}% (阈值:{thresholds['disk']}%)"
alerts.append(("磁盘告警", msg))
logger.warning(msg)
# 检查进程
try:
# 检查关键服务是否运行
result = subprocess.run(
["pgrep", "-a", "nginx"],
capture_output=True,
text=True
)
if result.returncode != 0:
msg = "Nginx服务未运行"
alerts.append(("服务告警", msg))
logger.error(msg)
except Exception as e:
logger.error(f"进程检查异常:{e}")
# 推送告警
for title, msg in alerts:
send_alert(title, msg, webhook_url)
return len(alerts) > 0
# ============= 定时监控 =============
def monitor_loop(thresholds, webhook_url, interval=60):
"""启动监控循环"""
logger.info("监控系统启动")
check_count = 0
alert_count = 0
try:
while True:
check_count += 1
if check_system_health(thresholds, webhook_url):
alert_count += 1
# 每小时打印一次统计
if check_count % 60 == 0:
logger.info(f"检查统计:总检查{check_count}次,告警{alert_count}次")
time.sleep(interval)
except KeyboardInterrupt:
logger.info("监控系统已停止")
except Exception as e:
logger.critical(f"监控系统异常退出:{e}")
# ============= 启动程序 =============
if __name__ == "__main__":
# 监控阈值配置
thresholds = {
"cpu": 80, # CPU超过80%告警
"memory": 85, # 内存超过85%告警
"disk": 90 # 磁盘超过90%告警
}
# 企业微信webhook地址(需要替换为实际地址)
webhook_url = "https://qyapi.weixin.qq.com/cgi-bin/webhook/send?key=xxx"
# 启动监控
monitor_loop(thresholds, webhook_url, interval=60)

这个脚本展示了如何整合:

  • psutil采集系统指标
  • subprocess执行系统命令检查进程
  • requests推送告警
  • logging记录所有操作和告警
  • 完整的异常处理和日志记录
Python第二章 运维核心工具库:高频场景专用模块
https://www.6ixblog.site/posts/linux-python-2/
作者
Licwic
发布于
2025-06-15
许可协议
CC BY-NC-SA 4.0