外观
15 - 监控与可观测性
概述
TDNS 提供完善的可观测性体系,包括 Prometheus 指标导出、结构化日志和运行时统计,帮助运维人员全面掌握服务器运行状态。
快速上手:三步开启监控
1. 配置
toml
# /etc/tdns/tdns.toml
[metrics]
enabled = true
listen = "127.0.0.1:9090" # 单个 IP:Port 字符串
path = "/metrics"
[server]
log-level = "info"
log-format = "json" # 结构化日志便于采集
[api]
enabled = true
listen = "127.0.0.1:8080"
auth-token = "your-secret-token"2. 启动 + 验证
bash
# 启动
sudo systemctl start tdns
# 查看指标
curl -s http://127.0.0.1:9090/metrics | head -20
# 健康检查
curl -s http://127.0.0.1:8080/api/v1/health
# 查看服务器状态
tdnsctl status3. 配置 Prometheus 抓取
yaml
# prometheus.yml
scrape_configs:
- job_name: 'tdns'
scrape_interval: 15s
static_configs:
- targets: ['localhost:9090']Prometheus 指标
端点
GET http://<server>:9090/metrics返回 Prometheus 文本格式(text exposition format),可直接被 Prometheus 抓取。
指标清单
| 指标 | 类型 | 标签 | 说明 |
|---|---|---|---|
tdns_queries_total | Counter | protocol, qtype, rcode | DNS 查询总数 |
tdns_query_duration_seconds | Histogram | protocol | 查询处理延迟 |
tdns_cache_hits_total | Counter | result | 缓存查找结果(hit/miss/negative_hit/stale_hit) |
tdns_cache_entries | Gauge | type (positive/negative) | 当前缓存条目数 |
tdns_upstream_queries_total | Counter | upstream, rcode | 上游查询总数 |
tdns_upstream_duration_seconds | Histogram | upstream | 上游查询延迟 |
tdns_rrl_dropped_total | Counter | reason (rate_limited/slipped) | RRL 丢弃的响应数 |
tdns_tsig_results_total | Counter | result (valid/invalid) | TSIG 验证结果 |
tdns_dnssec_validation_results_total | Counter | result | DNSSEC 验证结果(占位) |
tdns_zone_transfer_status | Gauge | - | 区域传输状态(占位) |
tdns_active_connections | Gauge | protocol | 当前活跃连接数 |
tdns_connections_total | Counter | protocol | 累计连接数 |
tdns_traffic_bytes_total | Counter | protocol, direction | 收发字节数 |
指标说明
查询指标
bash
curl -s http://localhost:9090/metrics | grep queries_totaltdns_queries_total{protocol="udp",qtype="A",rcode="NOERROR"} 123456
tdns_queries_total{protocol="tcp",qtype="AAAA",rcode="NXDOMAIN"} 789
tdns_queries_total{protocol="dot",qtype="A",rcode="NOERROR"} 456
tdns_queries_total{protocol="doh",qtype="A",rcode="NOERROR"} 123protocol:udp / tcp / dot / dohqtype:A / AAAA / CNAME / MX / TXT / NS / SOA / PTR / SRV / CAA / SVCB / HTTPS / ...rcode:NOERROR / NXDOMAIN / SERVFAIL / REFUSED / FORMERR / ...
查询延迟直方图
bash
curl -s http://localhost:9090/metrics | grep query_durationtdns_query_duration_seconds_bucket{protocol="udp",le="0.001"} 100000
tdns_query_duration_seconds_bucket{protocol="udp",le="0.005"} 115000
tdns_query_duration_seconds_bucket{protocol="udp",le="0.01"} 118000
...
tdns_query_duration_seconds_bucket{protocol="udp",le="+Inf"} 123456
tdns_query_duration_seconds_sum{protocol="udp"} 45.678
tdns_query_duration_seconds_count{protocol="udp"} 123456桶边界(秒):0.001, 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0, 10.0
缓存指标
bash
curl -s http://localhost:9090/metrics | grep cachetdns_cache_hits_total{result="hit"} 85234
tdns_cache_hits_total{result="miss"} 14766
tdns_cache_hits_total{result="negative_hit"} 3200
tdns_cache_hits_total{result="stale_hit"} 15
tdns_cache_entries{type="positive"} 12000
tdns_cache_entries{type="negative"} 3000上游指标
bash
curl -s http://localhost:9090/metrics | grep upstreamtdns_upstream_queries_total{upstream="8.8.8.8:53",rcode="NOERROR"} 12000
tdns_upstream_queries_total{upstream="8.8.8.8:53",rcode="FAILED"} 23
tdns_upstream_duration_seconds_bucket{upstream="8.8.8.8:53",le="0.05"} 11800
rcode="FAILED"表示上游超时或错误,可用于告警。
RRL 指标
bash
curl -s http://localhost:9090/metrics | grep rrltdns_rrl_dropped_total{reason="rate_limited"} 500
tdns_rrl_dropped_total{reason="slipped"} 50连接与流量指标
bash
curl -s http://localhost:9090/metrics | grep -E "connections|traffic"tdns_active_connections{protocol="dot"} 15
tdns_active_connections{protocol="doh"} 8
tdns_connections_total{protocol="dot"} 1500
tdns_connections_total{protocol="doh"} 800
tdns_traffic_bytes_total{protocol="udp",direction="rx"} 1048576
tdns_traffic_bytes_total{protocol="udp",direction="tx"} 2097152
tdns_traffic_bytes_total{protocol="dot",direction="rx"} 524288
tdns_traffic_bytes_total{protocol="dot",direction="tx"} 1048576Grafana 仪表板
推荐的关键监控面板:
| 面板 | 指标 | 查询 |
|---|---|---|
| QPS | tdns_queries_total | rate(tdns_queries_total[1m]) |
| 缓存命中率 | tdns_cache_hits_total | hit / (hit + miss) * 100 |
| 查询延迟 P99 | tdns_query_duration_seconds | histogram_quantile(0.99, ...) |
| 上游成功率 | tdns_upstream_queries_total | 1 - FAILED / total |
| RRL 拦截率 | tdns_rrl_dropped_total | rate(...[5m]) |
| 活跃连接数 | tdns_active_connections | 直接显示 |
| NXDOMAIN 比例 | tdns_queries_total | NXDOMAIN / total |
结构化日志
配置
toml
[server]
log-level = "info" # error/warn/info/debug/trace
log-format = "json" # json/text
# log-file = "/var/log/tdns/tdns.log" # 留空 = 输出到 stdout日志格式
JSON 格式(生产推荐):
json
{"timestamp":"2026-08-09T10:30:00Z","level":"INFO","target":"tdns::server","message":"DoT listener started","addr":"0.0.0.0:853"}文本格式(调试方便):
2026-08-09T10:30:00Z INFO tdns::server DoT listener started addr=0.0.0.0:853关键日志事件
| 事件 | 级别 | 说明 |
|---|---|---|
| 服务器启动 | INFO | 监听地址、模式、worker 数 |
| 监听器启动 | INFO | 各协议监听地址 |
| 区域加载 | INFO | 区域名称、记录数 |
| 热重载 | INFO | 重载的 Zone 数量 |
| ACL 拒绝 | DEBUG | 客户端 IP、拒绝原因 |
| 上游失败 | DEBUG | 上游地址、错误信息 |
| DoT/DoH 握手失败 | DEBUG | 客户端 IP、错误 |
| 连接超容量 | DEBUG | 客户端 IP |
| RRL 丢弃 | DEBUG | 客户端前缀、查询名 |
| 配置验证失败 | WARN | 验证错误详情 |
| 缓存淘汰 | DEBUG | 被淘汰的记录名 |
| 优雅关闭 | INFO | 停止原因、等待连接数 |
日志关键字段检索
bash
# 递归失败
journalctl -u tdns | grep "recursive.*failed\|SERVFAIL"
# 上游失败
journalctl -u tdns | grep "upstream.*failed\|timeout"
# ACL 拒绝
journalctl -u tdns | grep "REFUSED\|ACL"
# TLS 握手失败
journalctl -u tdns | grep "handshake.*failed"
# 连接超容量
journalctl -u tdns | grep "server at capacity"
# Zone 加载失败
journalctl -u tdns | grep "zone.*error\|zone.*failed"日志收集
推荐使用 vector / fluentd / filebeat 收集 JSON 日志:
yaml
# vector.toml
[sources.tdns_log]
type = "file"
include = ["/var/log/tdns/*.log"]
read_from = "beginning"
[sinks.elasticsearch]
type = "elasticsearch"
inputs = ["tdns_log"]运行时统计
通过 API 或 CLI 获取运行时统计:
bash
# 服务器状态
tdnsctl status
# 或
curl -s -H "Authorization: Bearer $TOKEN" \
http://127.0.0.1:8080/api/v1/server | jq
# 缓存统计
tdnsctl stats
# 或
curl -s -H "Authorization: Bearer $TOKEN" \
http://127.0.0.1:8080/api/v1/cache/stats | jq告警建议
| 告警条件 | 阈值 | 说明 |
|---|---|---|
| QPS 异常下降 | 低于基线 50% | 可能服务异常 |
| 缓存命中率下降 | < 50% | 缓存效率低,需调优 |
| 查询延迟 P99 | > 100ms | 响应慢 |
| 上游失败率 | > 5% | 上游 DNS 不稳定 |
| RRL 丢弃率突增 | > 100/min | 可能遭受攻击 |
| SERVFAIL 比例 | > 1% | 递归/转发异常 |
| 活跃连接数 | > 80% max | 连接数高 |
| NXDOMAIN 比例突增 | > 50% | 可能是扫描探测 |
| 服务不可达 | health check 失败 | 服务宕机 |
Prometheus 告警示例
yaml
groups:
- name: tdns
rules:
- alert: TdnsHighLatency
expr: histogram_quantile(0.99, rate(tdns_query_duration_seconds_bucket[5m])) > 0.1
for: 5m
labels:
severity: warning
annotations:
summary: "TDNS P99 latency > 100ms"
- alert: TdnsLowCacheHitRate
expr: |
rate(tdns_cache_hits_total{result="hit"}[5m]) /
(rate(tdns_cache_hits_total{result="hit"}[5m]) + rate(tdns_cache_hits_total{result="miss"}[5m])) < 0.5
for: 10m
labels:
severity: warning
annotations:
summary: "TDNS cache hit rate < 50%"
- alert: TdnsHighUpstreamFailRate
expr: |
rate(tdns_upstream_queries_total{rcode="FAILED"}[5m]) /
rate(tdns_upstream_queries_total[5m]) > 0.05
for: 5m
labels:
severity: critical
annotations:
summary: "TDNS upstream failure rate > 5%"