Nginx 配置与日志实战:从配置文件结构到 rewrite 规则

接手公司几台 Web 服务器时,我看到的第一样东西是:nginx.conf 被改过几十版,十几个站点全堆在一个文件里,日志统一写在一个 access.log。改一次配置要在一千多行里翻半天——那次改造是从”把配置理清楚”开始的。

一、安装:直接用官方源

CentOS 默认源里的 Nginx 版本偏旧,配置写法也落后。用官方仓库装:

1
cat /etc/yum.repos.d/nginx.repo
1
2
3
4
5
6
[nginx-stable]
name=nginx stable repo
baseurl=http://nginx.org/packages/centos/$releasever/$basearch/
gpgcheck=1
enabled=1
gpgkey=https://nginx.org/keys/nginx_signing.key
1
2
3
yum -y install nginx
nginx -v # nginx/1.24.0
systemctl start nginx && systemctl enable nginx

启停只有两种姿势,但同一时间只能用一种:交给 systemd(systemctl start|stop|reload nginx),或者用绝对路径自己管(/usr/sbin/nginx -s reload)。混用会出现”systemctl 显示没启动、进程却活着”的情况,团队里统一成一种就行。

二、配置结构:三层,业务全放 conf.d

Nginx 的配置是”总纲 + 分册”:主配置文件 nginx.conf 只放全局参数,业务站点的配置全部拆成独立文件放 conf.d/,主配置最后一行把它们 include 进来。

1
2
3
4
5
6
/etc/nginx/
├── nginx.conf # 主配置,只放全局参数
├── conf.d/ # 所有业务配置
│ ├── www.conf # 官网
│ └── admin.conf # 后台
└── ssl/ # 证书

nginx.conf 分三个块,记住各自的职责就行:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
# 全局块:管 Nginx 整体
user www; # 运行用户
worker_processes auto; # worker 数按 CPU 核数自动
error_log /var/log/nginx/error.log warn;
pid /run/nginx.pid;

# events 块:管连接能力
events {
use epoll; # Linux 下默认就是 epoll
worker_connections 10240; # 单 worker 最大连接数
}

# http 块:所有站点/代理的父容器
http {
include mime.types;
default_type application/octet-stream;
sendfile on;
tcp_nopush on;
keepalive_timeout 65;
gzip on;

include /etc/nginx/conf.d/*.conf; # 业务分册在这里进来
}

一个业务文件就是一个 server,用域名区分:

1
2
3
4
5
6
7
# /etc/nginx/conf.d/www.conf
server {
listen 80;
server_name www.example.com;
root /data/wwwroot/www;
index index.html;
}

三、location:请求路由的优先级

location 决定”这个 URL 用哪套配置处理”,是 Nginx 最核心也最容易写错的部分。优先级从高到低:

匹配符 规则 优先级
= 精确匹配 1
^~ 前缀匹配,命中后不再走正则 2
~ 正则,区分大小写 3
~* 正则,不区分大小写 4
/ 通用匹配 5

典型的生产写法:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
server {
listen 80;
server_name www.example.com;
root /data/www;

# 静态资源:本地直接返回,关日志、加缓存
location ~* \.(jpg|css|js)$ {
expires 7d;
access_log off;
}

# 后台接口:转发给后端 Java 服务
location /api/ {
proxy_pass http://127.0.0.1:8080;
}

# 后台管理页:限制内网访问
location /admin/ {
allow 192.168.10.0/24;
deny all;
}

location / {
index index.html;
}
}

root 和 alias 的区别这里必须点一下:location /download { root /package; } 实际找的是 /package/download;换成 alias /package; 才是去 /package 下拿。写错一个词,就是 404 和一小时排查的差距。

四、几个实用的基础模块

目录索引(当下载站用)

1
2
3
4
5
6
7
location /download {
alias /module;
autoindex on; # 列出目录文件
autoindex_exact_size off; # 显示 KB/MB,不显示精确字节
autoindex_localtime on; # 显示服务器时间
charset utf-8,gbk; # 中文文件名不乱码
}

状态监控(做连接数预警)

1
2
3
4
5
6
7
8
9
10
11
server {
listen 80;
server_name status.example.com;

location /nginx_status {
stub_status;
allow 172.18.1.0/24; # 只给内网看
allow 127.0.0.1;
deny all;
}
}

curl 一下能看到 Active connections、Reading/Writing/Waiting,接监控系统采集这两个值,连接数逼近上限时提前告警。

访问控制与限流

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
# 基于用户认证
location /nginx_status {
stub_status;
auth_basic "Auth";
auth_basic_user_file /etc/nginx/auth_conf;
}
# 生成密码文件(安装 httpd-tools 后)
htpasswd -b -c /etc/nginx/auth_conf ops 密码

# 连接限制:同一 IP 同时最多 1 个连接
http {
limit_conn_zone $remote_addr zone=conn_zone:10m;
}
server {
limit_conn conn_zone 1;
}

# 请求限制:每 IP 每秒 1 个请求,超出的进队列(burst),队列满返回 503
http {
limit_req_zone $binary_remote_addr zone=req_zone:10m rate=1r/s;
}
server {
limit_req zone=req_zone burst=3 nodelay;
limit_req_status 478; # 自定义状态码
error_page 478 /err.html; # 错误页换成自己的
}

限流的一个经验:别对 css/js/图片这类子资源做连接限制。一个页面动不动四十个静态请求,锁死连接数等于锁死正常用户;限流只针对 HTML 和接口。

五、rewrite:改名、跳转、伪静态

rewrite 用在三个场景:域名/协议跳转、URL 结构调整、伪静态。语法是 rewrite 正则 替换 [flag],flag 里最常用的是 redirect/permanent(返回 302/301)和 last/break(内部重写)。

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
# 案例1:/abc/1.html 实际访问 /ccc/bbb/2.html
location /abc {
rewrite (.*) /ccc/bbb/2.html redirect;
}

# 案例2:2018 年的目录结构改版,老链接自动跳新链接
location /2018 {
rewrite ^/2018/(.*)$ /2014/$1 redirect;
}

# 案例3:http 全站跳 https
server {
listen 80;
server_name www.example.com;
rewrite ^(.*) https://$server_name$1 redirect;
}

跳转规则写完必须先开 rewrite 日志验证,否则你根本不知道哪条规则命中了:

1
2
3
http {
rewrite_log on; # 日志级别 error_log ... notice
}

redirect(302)和 permanent(301)怎么选?链接做过改版、确定不会回退,用 301 让浏览器记住;只是临时调整(比如维护页),用小心的 302——301 会被浏览器缓存,用户端的”跳转记忆”清不掉,这个坑我遇到过:测试时用了 301,改回正确规则后自己电脑上一直跳旧地址,清了浏览器缓存才恢复。

六、日志:分站、切割、分析

日志拆分与格式

默认所有站点共用一个 access.log,出事时翻不动。规范做法是每个 server 单独一个日志文件:

1
2
3
4
5
6
7
8
9
10
server {
listen 80;
server_name www.example.com;
access_log /var/log/nginx/www.example.com.log main;

location /favicon.ico {
access_log off; # 这类请求没统计价值,直接关
return 200;
}
}

main 格式在 http 块里定义,生产上我会多带两个字段:

1
2
3
4
log_format main '$remote_addr - $remote_user [$time_local] "$request" '
'$status $body_bytes_sent "$http_referer" '
'"$http_user_agent" "$http_x_forwarded_for" '
'$request_time $upstream_response_time';

$request_time(总耗时)和 $upstream_response_time(后端耗时)是排查”慢在 Nginx 还是慢在后端”的钥匙,务必加上。

切割

Nginx 装完后 /etc/logrotate.d/nginx 已经配好,默认每天切割、保留 52 份、压缩归档,不用自己写脚本。唯一要记住的是:不要手动 rm 正在写的日志——进程还握着文件句柄,磁盘空间不会释放。要清空就用 > /var/log/nginx/access.log。

常用分析命令

1
2
3
4
5
6
7
8
# 访问量 TOP10 IP(抓爬虫/攻击源)
awk '{print $1}' /var/log/nginx/access.log | sort | uniq -c | sort -nr | head -10

# 状态码分布(看 404/502 的突增)
awk '{print $9}' /var/log/nginx/access.log | sort | uniq -c | sort -nr

# 慢接口 TOP(第 12 列是 $request_time)
awk '$NF > 1 {print $0}' /var/log/nginx/access.log | sort -k 12 -nr | head -10

需要图形化的时候上 goaccess,一条命令生成实时报表:

1
2
3
yum -y install goaccess
nohup goaccess -f /var/log/nginx/access.log -o /code/log/index.html \
-p /etc/goaccess/goaccess.conf --real-time-html &

七、变更规范

最后说一条纪律:改配置永远走”先 nginx -t 再 nginx -s reload“。-t 检查语法,reload 平滑重载——不断开任何连接。反例是直接 restart,业务高峰时等于主动制造一次故障。这个习惯看起来简单,但部署脚本里忘了加 nginx -t,一次缩进错误就足以让全站挂掉。