Ansible 流程控制实战:when、循环、handlers 与标签

写过几个剧本之后,你一定会遇到同一类需求:同一套剧本,web 机器该装 Nginx、db 机器该装 MySQL;配置文件只有真改了才重启服务;上百行的剧本想只跑其中一段调试。这些”分岔”和”联动”靠的就是流程控制。这篇把 when、循环、handlers、tags、include 和错误处理挨个过一遍。

一、when:一套剧本适配多主机

when 是条件和执行之间的开关,最典型的三个场景:

  • 不同系统装不同的包名(CentOS 用 yum,Ubuntu 用 apt);
  • 客户端和服务器共用一个剧本,但只有服务端需要推配置文件;
  • 源码编译先判断”是否已经装过”,避免重复执行。
1
2
3
4
5
6
7
8
9
10
11
12
13
- hosts: web_group
tasks:
- name: 安装 CentOS 的 httpd
yum:
name: httpd
state: present
when: ansible_facts['os_family'] == "RedHat"

- name: 安装 Debian 系的 apache2
apt:
name: apache2
state: present
when: ansible_facts['os_family'] == "Debian"

条件可以组合,也可以写成列表(列表里多个条件是”且”的关系):

1
2
3
4
5
6
7
8
9
10
# 括号分组:CentOS 6 或 Debian 7 才执行
- command: /sbin/shutdown -t now
when: (ansible_facts['distribution'] == "CentOS" and ansible_facts['distribution_major_version'] == "6") or
(ansible_facts['distribution'] == "Debian" and ansible_facts['distribution_major_version'] == "7")

# 列表写法:所有条件都满足才执行
- command: /sbin/shutdown -t now
when:
- ansible_facts['distribution'] == "CentOS"
- ansible_facts['distribution_major_version'] == "6"

还有一种基于返回值的匹配判断,match 是通配匹配、search 是包含匹配:

1
2
3
4
5
6
7
8
- name: 检查 nginx 配置
command: /usr/sbin/nginx -t
register: result
ignore_errors: yes

- name: 按 IP 段或主机名做条件
command: echo "命中"
when: ansible_facts['default_ipv4']['address'] is match "10.0.8.*"

is match 用的是通配符(*),is search 是子串包含——一个容易混的点:判断”输出里包含 successful”要用 search,用 match 必须从头匹配才成立。

二、一个真实剧本:rsync 的差异化部署

rsync 服务端和客户端常在一份剧本里,靠主机名分流:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
- hosts: rsync_all
tasks:
- name: 安装 rsync 软件包
yum:
name: rsync
state: present

- name: 创建 www 组
group:
name: www
gid: 666

- name: 仅备份服务器推送服务端配置
copy:
src: ./rsyncd.conf
dest: /etc/rsyncd.conf
mode: '0644'
when: ansible_hostname == "backup01"

- name: 仅备份服务器创建认证文件
copy:
content: 'rsync_backup:123'
dest: /etc/rsync.passwd
mode: '0600'
when: ansible_hostname == "backup01"

- name: 仅备份服务器启动服务
systemd:
name: rsyncd
state: started
enabled: yes
when: ansible_hostname == "backup01"

- name: 仅 web 服务器下发备份脚本
copy:
src: ./backup.sh
dest: /root/backup.sh
mode: '0755'
when: ansible_hostname is match "web*"

这就是 when 的核心价值:同一份剧本,撒到整个组上,每台机器各拿各的分支。不用为客户端和服务端各写一份,维护成本减半。

三、循环:重复的活只写一次

列表循环最简单:

1
2
3
4
5
6
7
8
9
10
- hosts: web_group
tasks:
- name: 安装多个软件包
yum:
name: "{{ item }}"
state: present
loop:
- httpd
- httpd-tools
- php

更常用的是字典循环,一次循环里带多组参数:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
- name: 批量创建用户(每组参数不同)
user:
name: "{{ item.name }}"
groups: "{{ item.groups }}"
state: present
loop:
- { name: 'ops', groups: 'ops' }
- { name: 'bizdev', groups: 'dev' }

- name: 拷贝多个文件(源、目标、权限都不同)
copy:
src: "{{ item.src }}"
dest: "{{ item.dest }}"
mode: "{{ item.mode }}"
loop:
- { src: "./httpd.conf", dest: "/etc/httpd/conf/", mode: "0644" }
- { src: "./upload.php", dest: "/var/www/html/", mode: "0600" }

loop 是新写法,老笔记里常见的 with_items 依然可用(两者在简单列表场景等价)。新写剧本建议统一用 loop,语义更清楚。

四、notify 与 handlers:配置变了才重启

这是 Ansible 里设计最精妙的一对概念。问题很直接:改了 nginx 配置要重启服务,但总不能每次执行剧本都重启一次吧?

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
tasks:
- name: 安装 httpd
yum:
name: httpd
state: present

- name: 分发 httpd 配置
template:
src: ./httpd.j2
dest: /etc/httpd/conf/httpd.conf
notify: 重启httpd # 只有这个任务真的改了文件才会触发

- name: 启动 httpd
service:
name: httpd
state: started
enabled: yes

handlers:
- name: 重启httpd
systemd:
name: httpd
state: restarted

机制是这样的:配置模块执行后如果 changed=true(内容真的变了),就触发 notify 指定的 handler;所有 task 执行完,统一执行被触发的 handlers;同一个 handler 被多次 notify 也只执行一次。配置没变,handler 完全不执行,服务不受打扰。

五条行为规则要记住,都是排查”明明改了配置为什么没重启”的钥匙:

  1. 无论多少个任务 notify 了同一个 handler,它只会在所有 task 结束后运行一次;
  2. handler 所在的任务如果因为条件判断没执行,handler 也不会执行;
  3. handler 在该 play 末尾运行;想在中间就执行,用 meta: flush_handlers;
  4. play 中途失败,后面的 handler 不会执行(除非 --force-handlers);
  5. 不能把 handler 当普通 task 用——它只响应 notify。

五、标签:大剧本的分段执行

剧本一长,调试就成了问题:只想跑”配置更新”这一段,不想重装软件。给任务打标签解决:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
- hosts: web_group
tasks:
- name: 安装 httpd
yum:
name: httpd
state: present
tags: install_httpd

- name: 分发配置
template:
src: ./httpd.j2
dest: /etc/httpd/conf/httpd.conf
notify: 重启httpd
tags: [config_httpd, httpd_server]

- name: 启动服务
service:
name: httpd
state: started
tags: service_httpd
1
2
3
4
ansible-playbook tag.yml --list-tags              # 先看有哪些标签
ansible-playbook tag.yml -t config_httpd # 只跑打这个标签的任务
ansible-playbook tag.yml -t install_httpd,config_httpd
ansible-playbook tag.yml --skip-tags service_httpd # 跳过指定标签

一个实用的团队约定:给每个任务都打”模块标签 + 业务标签”两个,比如 [config_httpd, httpd_server]——前者按动作执行,后者按业务整体执行,两种切法都支持。

六、文件复用:include_tasks 与 import_playbook

一个剧本上百行就该拆了。两种复用方式:

1
2
3
4
5
6
# include_tasks:引用任务文件(动态,运行时展开)
- hosts: web_group
tasks:
- include_tasks: task_install.yml
- include_tasks: task_configure.yml
- include_tasks: task_start.yml
1
2
3
4
5
# import_playbook:拼装多个完整剧本,做成"一键入口"
# site.yml
- import_playbook: httpd.yml
- import_playbook: nfs.yml
- import_playbook: rsync.yml

区别在加载时机:include_tasks 是运行时动态加载,import_playbook 是开始前静态引入。日常拆分任务用前者,做统一入口用后者。有了入口文件,一套架构的部署就是一条命令:

1
ansible-playbook site.yml

七、错误处理:三个实用开关

ignore_errors:默认剧本遇到失败会立刻停下,但有些任务”失败也不影响继续”(比如检测类的命令):

1
2
3
4
5
6
7
8
- name: 检测命令(失败也继续)
command: /bin/false
ignore_errors: yes

- name: 后续任务照常执行
file:
path: /tmp/task_done.txt
state: touch

force_handlers:play 中途失败时,已触发的 handler 默认不执行——这可能导致”配置改了但服务没重载”的中间状态。加这个开关强制把 handler 跑完:

1
2
- hosts: web_group
force_handlers: yes

changed_when:有些命令(比如检测类 shell)其实没改任何东西,但 Ansible 认为它 always changed。这两个作用:一是输出更准,二是避免误触发 handler:

1
2
3
4
5
6
7
8
9
10
11
- name: 检查 httpd 状态(抑制 changed)
shell: netstat -lntup | grep httpd
register: check_httpd
changed_when: false

- name: 检查配置语法(按输出判断)
shell: /usr/sbin/httpd -t
register: httpd_check
changed_when:
- httpd_check.stdout.find('OK')
- false

一句总结

流程控制这几个工具,解决的是同一件事:让剧本带上判断力。when 决定”该不该做”、loop 决定”做几遍”、handlers 决定”什么时候做”(配置真的变了才做)、tags 决定”这次只做哪部分”。把它们用顺,剧本才会从一条直线变成能适应复杂环境的程序。