Ansible 常用模块实战:从软件包到数据库
Ansible 用起来之后,我给自己整理过一份”模块速查表”——不是背参数,而是记住”哪类活对应哪个模块、关键参数是哪几个”。这份表后来成了我写剧本时的第一参考,这篇把它整理出来,从最常用的软件包、文件、服务,一直到用 community.mysql 批量管理数据库。
一、命令执行三兄弟
1 2 3 4 5 6 7 8 9 10 11
| - name: 查看主机名 command: hostname
- name: 查 nginx 进程 shell: ps -ef | grep nginx
- name: 执行本地检查脚本 script: /opt/ops-ansible/scripts/check.sh
|
原则很简单:能用 command 就用 command(更安全,没有 shell 注入问题),需要管道才换 shell。
二、软件包管理
1 2 3 4 5 6 7
| - name: 安装指定版本 Nginx yum: name: nginx-1.20.1 state: present enablerepo: epel update_cache: yes
|
state 的取值要分清:present/installed 是安装、absent/removed 是卸载、latest 是升到最新。批量管控时用固定版本号,别用 latest——“所有机器版本一致”比”都是最新的”重要。
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19
| - name: 配置内网 Base 源 yum_repository: name: local-base description: 内网 CentOS7 Base 源 baseurl: http://172.18.1.100/yum/centos7/base gpgcheck: no enabled: yes file: internal-yum
- name: 上传并安装监控 Agent copy: src: ../files/monitor-agent.rpm dest: /tmp/monitor-agent.rpm - name: 本地安装 rpm rpm: name: /tmp/monitor-agent.rpm state: present
|
三、文件管理:分发与配置修改
1 2 3 4 5 6 7 8 9
| - name: 分发 rsyncd 主配置 copy: src: ../files/rsyncd.conf dest: /etc/rsyncd.conf owner: root group: root mode: '0644' backup: yes
|
copy 的 backup: yes 在生产上必开:它会把被覆盖的旧文件按时间戳备份,出问题能马上找回原文件。
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21
| - name: 创建网站根目录 file: path: /data/www/html state: directory owner: nginx group: nginx mode: '0755' recurse: yes
- name: 创建应用软链 file: src: /opt/app-v1.2.3 dest: /opt/app-current state: link
- name: 清理临时目录 file: path: /tmp/install_pkg state: absent
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
| - name: 下载 Node Exporter get_url: url: http://172.18.1.100/software/node_exporter-1.8.2.linux-amd64.tar.gz dest: /tmp/node_exporter.tar.gz checksum: sha256:<校验值> force: no
- name: 解压到安装目录 unarchive: src: /tmp/node_exporter.tar.gz dest: /opt/node_exporter remote_src: yes extra_opts: [--strip-components=1] creates: /opt/node_exporter/node_exporter
|
creates 参数是幂等的关键:它让”这个文件已存在就跳过”,重复执行剧本不会报错也不会重复解压。
改配置文件有三件套,用哪个看场景:
| 模块 |
适用场景 |
特点 |
| lineinfile |
改单行配置,确保某行存在/不存在 |
按整行匹配,天然幂等 |
| replace |
全局字符串替换,多处同类修改 |
只替换匹配片段,不关心整行 |
| blockinfile |
插入/更新一段多行配置 |
用标记行包裹,不会重复插入 |
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35
| - name: 关闭 SELinux lineinfile: path: /etc/selinux/config regexp: '^SELINUX=' line: 'SELINUX=disabled' backup: yes
- name: 禁止 root 直登 lineinfile: path: /etc/ssh/sshd_config regexp: '^PermitRootLogin' line: 'PermitRootLogin no' backup: yes notify: 重启sshd
- name: 批量替换监听地址 replace: path: /etc/redis/redis.conf regexp: '127\.0\.0\.1' replace: '0.0.0.0' backup: yes
- name: 添加内核优化参数 blockinfile: path: /etc/sysctl.conf block: | net.ipv4.tcp_syncookies = 1 net.ipv4.tcp_max_syn_backlog = 8192 fs.file-max = 655350 marker: "# {mark} KERNEL OPTIMIZE BLOCK" backup: yes
|
四、服务管理:service 与 systemd
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
| - name: 启动 rsyncd 并设为开机自启 service: name: rsyncd state: started enabled: yes
- name: 分发 node_exporter 单元文件并重启 copy: src: ../files/node_exporter.service dest: /usr/lib/systemd/system/node_exporter.service notify: 重启node_exporter
- name: 重启node_exporter systemd: name: node_exporter state: restarted enabled: yes daemon_reload: yes
|
state 取值里有个容易忽略的点:restarted 每次执行都会真的重启,不幂等;started 才是”没启动才启动”。日常用 started,只有配置变更触发时才在 handler 里用 restarted。
五、用户、定时任务与挂载
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24
| - name: 批量创建业务用户 user: name: "{{ item.name }}" group: biz shell: /bin/bash create_home: yes loop: - { name: 'biz_user1' } - { name: 'biz_user2' }
- name: 创建临时外包账号 user: name: temp_outsource shell: /bin/bash expires: "{{ '2024-12-31' | to_datetime('%Y-%m-%d') | int }}"
- name: 推送运维公钥 authorized_key: user: opsadmin key: "{{ lookup('file', '/root/.ssh/id_rsa.pub') }}" state: present
|
expires 参数很适合外包/临时账号管理——比”到期记得手动删”靠谱得多。
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17
| - name: 每天凌晨 2 点 MySQL 全量备份 cron: name: "mysql-full-backup-daily" user: root minute: "0" hour: "2" job: "/usr/bin/mysqldump -uroot -p'Root@123456' --all-databases > /backup/mysql_full_$(date +\\%Y\\%m\\%d).sql 2>> /var/log/mysql_backup.log" backup: yes
- name: 工作时段健康巡检(每 10 分钟) cron: name: "app-health-check-worktime" minute: "*/10" hour: "9-18" weekday: "1-5" job: "/opt/scripts/health_check.sh >> /var/log/health_check.log 2>&1"
|
注意 cron 里 backup: yes 同样建议开着,批量改 crontab 出问题时能对照原文件。备份任务里的密码明文是个反面教材——生产应该走 vault 或者独立的密码文件,后面讲变量的文章会展开。
1 2 3 4 5 6 7 8 9
| - name: 挂载数据盘 mount: path: /data src: /dev/mapper/data_vg-data_lv fstype: xfs opts: defaults,noatime state: mounted backup: yes
|
state 的四种取值值得记:mounted(挂载且写 fstab)、present(只写 fstab)、unmounted(卸载但保留 fstab)、absent(卸载并清 fstab)。
六、防火墙与主机信息
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
| - name: 批量放行常用端口 ansible.posix.firewalld: port: "{{ item.port }}/{{ item.proto }}" permanent: yes immediate: yes state: enabled loop: - { port: "80", proto: "tcp" } - { port: "443", proto: "tcp" } - { port: "3306", proto: "tcp" }
- name: 限制数据库端口来源 ansible.posix.firewalld: rich_rule: 'rule family="ipv4" source address="172.18.1.0/24" port port="3306" protocol="tcp" accept' permanent: yes immediate: yes state: enabled zone: internal
|
permanent 和 immediate 建议都写 yes:前者保证重启后仍生效,后者保证当前立刻生效。只写 permanent 的话,端口要等下次重启才通,这个坑我踩过一次——放行完端口,服务还是连不上,排查半天。
主机信息采集用 setup 模块(playbook 默认自动执行,即 gather_facts),采集结果放在 ansible_facts 里,巡检类剧本全靠它:
1 2 3 4 5 6
| - debug: msg: "主机名:{{ ansible_facts['hostname'] }}" - debug: msg: "内存:{{ ansible_facts['memtotal_mb'] }} MB,CPU:{{ ansible_facts['processor_vcpus'] }} 核" - debug: var: ansible_facts['mounts']
|
数据库这块用官方集合 community.mysql,先装:
1 2 3 4 5
| ansible-galaxy collection install community.mysql
yum install -y python3-PyMySQL
|
认证安全有两条路:被控端在 /root/.my.cnf 配好免密登录(模块自动读取,剧本里不出现密码);或者用 ansible-vault 加密的变量传入。密码直接写在剧本里是不行的,剧本是进版本库的。
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
| - name: 创建订单业务库 community.mysql.mysql_db: name: app_order state: present encoding: utf8mb4 collation: utf8mb4_unicode_ci
- name: 备份订单库(gzip 压缩、不锁表) community.mysql.mysql_db: name: app_order state: dump target: /data/backup/app_order_2024-06-23.sql.gz single_transaction: yes quick: yes
|
1 2 3 4 5 6 7 8 9 10
| - name: 创建业务只读账号 community.mysql.mysql_user: name: app_ro host: '10.0.8.%' password: 'StrongPass@2024' priv: 'app_order.*:SELECT' append_privs: yes update_password: on_create state: present
|
账号管理有三条生产规范:host 精确到网段,不写 %;append_privs: yes 避免覆盖已有权限;update_password: on_create 让”同步账号”和”改密码”两件事解耦。
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22
| - name: 连接数打满时临时调大 community.mysql.mysql_variables: variable: max_connections value: 2000 mode: global
- name: 采集实例基础信息 community.mysql.mysql_info: filter: [version, databases, status] register: mysql_info_result
- name: 获取从库复制状态 community.mysql.mysql_replication: mode: getslave register: slave_status
- name: 输出同步延迟 debug: msg: "延迟:{{ slave_status.Seconds_Behind_Master }} 秒"
|
mysql_replication 配合 when 断言可以做自动告警:延迟超过阈值就 fail 掉这一台,巡检结果里直接能看出来。最后是 mysql_query,用来执行模块覆盖不到的自定义 SQL——生产上有个纪律:执行 DML/DDL 前必须备份,禁止执行没有 WHERE 条件的 UPDATE/DELETE。
1 2 3 4
| - name: 清理半年前的访问日志 community.mysql.mysql_query: login_db: app_order query: "DELETE FROM access_log WHERE create_time < DATE_SUB(NOW(), INTERVAL 6 MONTH)"
|
两句经验
模块不需要背全,抓住几条主线就行:安装用 yum、分发用 copy、改配置用 lineinfile 三件套、服务联动靠 notify/handlers、数据库操作走 community.mysql。真正的功夫在参数的”生产值”上——backup 开不开、version 固定不固定、权限给多少,这些决定了剧本能不能在生产放心地跑第二遍。