// 红队渗透 · 2026-08-15

dirsearch v0.5.0 字典扫描进阶玩法

概述

dirsearch v0.5.0(2026-08-14 发布)是原 Python 版作者 maurosoria 的全面重构版本,提供三种运行后端:

后端 命令 特点
Python asyncio 默认 跨平台,功能完整
Python 线程 --sync 兼容旧版行为
Rust native --request-backend native 单文件二进制,~50MB,无 Python 依赖,性能更高

同时字典生成也支持 --wordlist-backend native 用 Rust 引擎生成,不受 Python GIL 和内存限制。


一、字典系统深度解析

1.1 三层字典体系

v0.5.0 将字典拆为三层,不再是单一 flat 文件:

第一层:分类字典(db/categories/)

按技术栈拆分为 32 个文件,配合 --wordlist-categories 按需加载:

通用类: common, web, conf, backups, db, logs, keys, vcs, extensions
PHP:    laravel, wordpress, codeigniter, symfony, yii, cakephp, joomla, drupal, magento
Python: django, flask, fastapi
Java:   jsp, jsf, spring
.NET:   aspx, mvc, core
Node:   express
CF:     coldfusion
基础设施: docker, k8s, aws
# 只扫 PHP 目标
dirsearch -u https://target.com --wordlist-categories php

# 全量
dirsearch -u https://target.com --wordlist-categories all

# 组合
dirsearch -u https://target.com --wordlist-categories php,dotnet,api,infra

第二层:模板字典(db/templates/)

用占位符一行生成数十条路径,内置模板文件:

模板文件 占位符 用途
admin.txt %ADMIN_OP%, %SUBJECT% 后台管理路径
api.txt %API_VERSION%, %CRUD_OP%, %SUBJECT% API 端点
auth.txt %AUTH_OP% 认证相关路径
crud.txt %CRUD_OP%, %SUBJECT% CRUD 操作路径
backups.txt %DATE%, %EXT% 备份文件
logs.txt %ENV%, %DATE% 日志文件
db.txt %ENV%, %SUBJECT% 数据库相关

可用占位符一览:

占位符 展开内容示例
%SUBJECT% users, products, orders, payments, invoices, reports, categories, comments, messages, notifications, settings, profiles, articles, images, files, logs, pages, posts, tags, items
%CRUD_OP% create, read, update, delete, add, edit, remove, view, list, new, manage, save, delete, modify, show, index, destroy, patch, put, post
%AUTH_OP% login, logout, register, signin, signup, forgot, reset, verify, authenticate, authorize, token, session, password, callback, sso, oauth, openid, saml
%ADMIN_OP% dashboard, settings, users, roles, permissions, config, logs, backup, restore, report, analytics, audit, queue, cache, cron, schedule, health, status, monitor, debug
%ENV% dev, staging, test, prod, production, development, uat, qa, sandbox, local, live, demo, beta, alpha, release
%DATE% 2024, 2025, 2024-01, 2025-01-01, january, jan, 01, 2024-01-15, 01012024, 2024_01_15
%API_VERSION% v1, v2, v3, v1.0, v2.0, api, latest, beta, v1_0, v1-0, 1.0, 2.0, 3.0
%CATEGORY:name% 引用 db/categories/ 中的分类字典内容
%EXT% 由 -e 参数指定的扩展名

示例:admin.txt 模板展开

模板内容:
  %ADMIN_OP%.%EXT%
  admin/%ADMIN_OP%
  %ADMIN_OP%/%SUBJECT%

假设 -e php,展开后:
  dashboard.php, settings.php, users.php
  admin/dashboard, admin/settings, admin/users
  dashboard/users, settings/users, dashboard/reports

第三层:自定义字典

# 混合内置分类 + 自定义字典
dirsearch -u https://target.com \
  --wordlist-categories php,java \
  -w /path/to/your/custom.txt,/path/to/seclists/Discovery/Web-Content/*.txt

# 预览字典条目数(不扫描)
dirsearch --wordlist-categories all --wordlist-status

# 限制字典大小防 OOM
dirsearch --wordlist-max-size 100000

1.2 字典变换技巧

# 后缀追加
dirsearch -e php -u https://target --suffixes ~

# 前缀追加
dirsearch -e php -u https://target --prefixes .,admin,_
# 对 tools 生成: .tools, admintools, _tools

# 大小写变换
dirsearch -U -u https://target  # 全大写
dirsearch -L -u https://target  # 全小写
dirsearch -C -u https://target  # 首字母大写

# 强制扩展名追加(对不带 %EXT% 的字典)
dirsearch -e php,asp -f -u https://target
# 对 admin 生成: admin, admin.php, admin.asp, admin/

# 覆盖已有扩展名
dirsearch -e jsp,jspa --overwrite-extensions -u https://target
# 对 login.html 生成: login.html, login.jsp, login.jspa

# 排除特定扩展名
dirsearch -e php --exclude-extensions jsp -u https://target

二、高级过滤系统

2.1 匹配/过滤矩阵

v0.5.0 提供 8 维过滤维度,每维可分别做“匹配(只保留符合的)“和”过滤(排除符合的)“:

维度 匹配参数 过滤参数
状态码 --mc --fc
响应体大小 --ms --fs
单词数 --mw --fw
行数 --ml --fl
响应体正则 --mr --fr
响应头文本 --match-header --filter-header
响应头正则 --match-header-regex --filter-header-regex
响应时间(ms) --mt --ft
# 只保留 200 和 403
dirsearch -u https://target --mc 200,403

# 排除 301,302,500-599
dirsearch -u https://target --fc 301,302,500-599

# 按大小过滤
dirsearch -u https://target --fs 0,4KB,10KB

# 按响应时间过滤(排除响应慢的)
dirsearch -u https://target --ft '>500'

# 正则匹配
dirsearch -u https://target --mr 'admin|dashboard|config'

# 排除包含特定文本的响应
dirsearch -u https://target --exclude-text 'Not Found'

2.2 逻辑组合:and / or

默认多个条件是 or 关系,只要匹配任一即可。可以用 --mmode 和 --fmode 切换为 and:

# 同时满足状态码 200 且响应体包含 admin
dirsearch -u https://target \
  --mc 200 --mr 'admin|dashboard' \
  --mmode and

# 排除同时满足 301 和空响应体的
dirsearch -u https://target \
  --fc 301 --fs 0 \
  --fmode and

2.3 通配符过滤与去重

# 强制通配符校准
dirsearch -u https://target --auto-calibration

# 阈值过滤:相同内容的响应出现 N 次后自动过滤
dirsearch -u https://target --filter-threshold 3

2.4 基于参考页面的过滤

# 排除与 404 页面相似的响应
dirsearch -u https://target --exclude-response /404.html

三、智能扫描策略

3.1 递归扫描

# 基础递归
dirsearch -u https://target.com -r

# 深度递归(每层目录都扫)
dirsearch -u https://target.com --deep-recursive -R 3

# 强制递归(非目录路径也递归)
dirsearch -u https://target.com --force-recursive

# 仅对特定状态码递归
dirsearch -u https://target.com --recursion-status 200,301,403

3.2 爬虫模式

# 从响应中爬取新路径自动加入扫描
dirsearch -u https://target.com --crawl

3.3 扫描控制

# 总扫描时长限制
dirsearch -u https://target --max-time 300

# 单目标时长限制
dirsearch -u https://target --target-max-time 60

# 速率限制
dirsearch -u https://target --max-rate 100

# 请求延迟
dirsearch -u https://target --delay 0.5

# 遇到特定状态码跳过目标
dirsearch -u https://target --skip-on-status 403,404

# 出错即退出
dirsearch -u https://target --exit-on-error

3.4 请求控制

# 自定义请求方法+数据
dirsearch -u https://target.com/api \
  -m POST \
  -d '{"id":1}' \
  -H 'Content-Type: application/json'

# 从文件加载原始 HTTP 请求
dirsearch --raw request.txt --scheme https

# 认证
dirsearch -u https://target.com \
  --auth admin:password123 \
  --auth-type basic

dirsearch -u https://target.com \
  --auth 'eyJhbGciOiJI...' \
  --auth-type bearer

# 客户端证书
dirsearch -u https://target.com \
  --cert-file client.crt \
  --key-file client.key

# Cookie
dirsearch -u https://target.com --cookie 'session=abc123'

# User-Agent 随机化
dirsearch -u https://target.com --random-agent

# 跟随重定向
dirsearch -u https://target.com -F

# 显示重定向历史
dirsearch -u https://target.com --redirects-history

四、代理与网络

4.1 多代理轮换

# 多代理轮换
dirsearch -u https://target.com \
  -p http://proxy1:8080 \
  -p http://proxy2:8080 \
  -p socks5://proxy3:1080

# 从文件加载代理列表
dirsearch -u https://target.com --proxies-file proxies.txt

# 代理认证
dirsearch -u https://target.com -p http://proxy:8080 \
  --proxy-auth user:pass

# 重放代理(路径走代理,正常请求不走)
dirsearch -u https://target.com --replay-proxy http://replay:8080

4.2 网络配置

# 指定网卡
dirsearch -u https://target.com --interface eth0

# 指定目标 IP(省去 DNS 解析)
dirsearch -u https://target.com --ip 10.10.10.10

# 超时与重试
dirsearch -u https://target.com --timeout 10 --retries 3

五、Nmap 联动

# 先扫 Web 端口
nmap -sV -p 80,443,8080,8443,9000,5000,3000 -oX web.xml <target>

# 自动导入
dirsearch --nmap-report web.xml

# 配合 CIDR 批量
dirsearch --cidr 10.10.10.0/24
dirsearch -l urls.txt
dirsearch --stdin < urls.txt

六、报告与输出

6.1 多格式同时输出

dirsearch -u https://target.com \
  -O json,xml,md,csv,html,sqlite \
  -o result

# 自动命名
dirsearch -u https://target.com -o result
# 生成 result.json, result.xml, result.md 等

6.2 数据库直写

# 写 MySQL
dirsearch -u https://target.com \
  -O mysql \
  --mysql-url mysql://user:pass@localhost:3306/dirsearch

# 写 PostgreSQL
dirsearch -u https://target.com \
  -O postgres \
  --postgres-url postgres://user:pass@localhost:5432/dirsearch

6.3 日志与静默模式

# 详细日志
dirsearch -u https://target.com --log scan.log

# 静默模式(只输出结果)
dirsearch -u https://target.com -q

# 完全禁用 CLI 输出
dirsearch -u https://target.com --disable-cli

# 显示详细信息(响应时间 + Content-Type)
dirsearch -u https://target.com -v

七、会话与断点续扫

# 自动保存会话(Ctrl+C 暂停)
dirsearch -u https://target.com -s session.json

# 列出所有可恢复会话
dirsearch --list-sessions

# 按 ID 恢复
dirsearch --session-id 3

# 会话路径:默认 ~/.dirsearch/sessions/(bundle 版)
# 或运行目录下的 sessions/

八、Python API 模式

可将 dirsearch 当库导入:

from dirsearch.core import FuzzerConfig, DirsearchFuzzer

config = FuzzerConfig(
    url="https://target.com",
    extensions=["php", "asp"],
    wordlist_categories=["php", "common"],
)

fuzzer = DirsearchFuzzer(config)
results = fuzzer.run()

for result in results:
    print(f"{result.status} {result.path} ({result.size}b)")

九、实战场景组合

场景 1:快速扫描 PHP 站点

dirsearch -u https://target.com \
  --request-backend native \
  --wordlist-backend native \
  --wordlist-categories php,common,web,backups \
  -e php \
  --mc 200,403,301 \
  --fc 500-599 \
  --random-agent \
  --max-rate 200

场景 2:深度 API 渗透

dirsearch -u https://api.target.com \
  -w db/templates/api.txt \
  -e json,xml \
  -m POST \
  -H "Authorization: Bearer <token>" \
  -H "Content-Type: application/json" \
  -d '{"id":1}' \
  --mc 200,201,403 \
  --mr 'error|unauthorized|admin|flag' \
  --deep-recursive -R 3 \
  --crawl

场景 3:后台管理暴力探测

dirsearch -u https://target.com/admin \
  -w db/templates/admin.txt,db/templates/auth.txt \
  -e php,html \
  --prefixes .,admin,_ \
  --suffixes ~,.bak,.old,.swp \
  --mc 200,403,301 \
  --mr 'dashboard|login|settings|config' \
  --filter-threshold 5

场景 4:批量目标扫描

# 1. Nmap 扫内网
nmap -sV -p 80,443,8080 -oX web.xml 10.10.10.0/24

# 2. 导入 dirsearch
dirsearch --nmap-report web.xml \
  --request-backend native \
  --wordlist-categories all \
  -e php,asp,jsp \
  --target-max-time 120 \
  -O json,csv \
  -o batch_result

场景 5:WordPress 专项

dirsearch -u https://target.com \
  --wordlist-categories php/wordpress \
  -e php \
  --crawl \
  -r \
  --mc 200,301,403 \
  --mr 'wp-admin|wp-content|wp-includes|wp-config|xmlrpc'

场景 6:利用模板字典做框架指纹后的精准扫描

# 识别出是 Laravel + Vue 后
dirsearch -u https://target.com \
  --wordlist-categories php/laravel,node/express \
  -w db/templates/api.txt,db/templates/crud.txt \
  -e php,json \
  --mc 200,201,403,405 \
  --deep-recursive

十、Rust Native 版本说明

优势

  • 单文件 ~50MB,无 Python 依赖
  • 字典生成 Rust 版不受 Python 内存限制
  • HTTP 请求性能更高(连接池、复用)

局限

  • 字典和配置都打包在二进制内,运行时释放到 /tmp/_MEI*,进程退出自动清理
  • 要修改内置字典需用 -w 指定外部文件覆盖
  • 某些功能可能依赖 Python 后端(如 --sync 线程模式)

下载

wget -O /usr/local/bin/dirsearch \
  https://github.com/maurosoria/dirsearch/releases/download/v0.5.0/dirsearch-v0.5.0-linux-x64-native-rust
chmod +x /usr/local/bin/dirsearch

总结

dirsearch v0.5.0 不再只是一个“扫目录的工具”,而是一个完整的 Web 内容发现引擎:

  • 三层字典:分类字典精准选型、模板字典批量裂变、自定义字典灵活扩展
  • 八维过滤:状态码/大小/单词数/行数/正则/响应头/时间,支持 and/or 组合
  • 三套引擎:async/threaded/native-rust,适应不同场景
  • 全流程覆盖:Nmap 联动导入 → 智能扫描 → 会话续扫 → 多格式报告

用好这些功能,字典扫描这件事可以做到“每一发请求都不浪费”。

原文 https://blog.csdn.net/2301_79518550/article/details/163782534