Python后端爬虫专题13:别反复下载没变化的详情——ETag、Last-Modified与增量采集
发布时间:2026/9/28 8:08:34来源:尧图网络
Python后端爬虫专题13别反复下载没变化的详情——ETag、Last-Modified与增量采集上一篇练习完整答案完整离线重放脚本如下。它先使用真实 Fixture 调用快照模块再从返回标识读回字节最后交给共享解析器断言失败时能明确区分快照损坏与选择器损坏。frompathlibimportPathfromtempfileimportTemporaryDirectoryfromjobradar.parsingimportparse_detailfromjobradar.snapshotsimportFileSnapshotStore fixturePath(tests/fixtures/job-detail.html).read_bytes()withTemporaryDirectory()asdirectory:storeFileSnapshotStore(Path(directory))refstore.save(http://target.test/jobs/python-backend-001,fixture,{content-type:text/html; charsetutf-8,etag:v1},)savedstore.load(ref.snapshot_id)jobparse_detail(saved.body.decode(utf-8),saved.metadata.url)assertjob.titlePython 后端工程师相同正文的 SHA-256 相同因此两次保存得到同一 snapshot_id正文哪怕只变化一个字节哈希也应改变。可保存的元数据包括 Content-Type、ETag、Last-Modified、最终 URL、大小和抓取时间Cookie、Set-Cookie、Authorization、代理凭据必须排除。会话 Cookie 泄露后拿到快照目录的人可能冒用登录状态访问内部系统风险远大于一次职位数据泄露。先算一笔无效下载账一家公司有 20 万个详情页每页 60 KB每小时重跑一次。如果职位几天才变化却每次下载全部正文一天流量约为 288 GB还会给对方服务制造不必要压力。增量采集的第一步不是上 Kafka而是尊重 HTTP 已有的缓存验证机制。服务器响应详情时给出ETag: abcJobRadar 将它与职位记录一起保存。下次请求发送If-None-Match: abc内容不变时服务器返回 304没有响应正文变化时返回 200、新正文和新 ETag。Last-Modified/If-Modified-Since是时间版本的同类机制但时间精度和代理行为可能导致误差服务器同时支持时通常优先 ETag。三次采集在流水线里发生了什么第一次没有 validators详情返回 200。流水线保存快照、解析、校验并 upsert报告created3。第二次 Repository 按(tenant_id, source_url)读出 ETagHttpFetcher 设置条件头TargetLab 返回三个 304流水线只增加not_modified3不解析空正文、不写新快照。第三次我们修改一个职位描述只有它返回 200 并触发updated1另外两个仍为 304。运行贯穿 HTTP、存储、解析与数据库的检查点cd project.\.venv\Scripts\python.exe-m pytest tests\test_targetlab_journey.py::test_real_targetlab_journey_creates_skips_and_updates_jobs-q这个测试使用进程内真实 ASGI TargetLab不伪造流水线返回值。它最后查询数据库确认只有三行而且被更新行真的含新描述。若只是断言请求头存在仍无法证明 304 分支没有误删数据。304 不是“抓取失败”304 的意义是“你给出的本地版本仍然有效”不是没有职位。FetchResult因而保留not_modified标志调用者看到它就直接计数。若把所有非 200 都交给raise_for_status304 会被错误处理若把 304 当空 HTML 解析又会得到“缺少 h1”的假告警。需要注意列表页目前仍正常下载。因为列表页决定新职位发现和分页如果完全跳过它就无法知道新增入口。更大的系统会对列表也做条件请求并安排低频全量巡检防止缓存头配置错误让新数据长期不可见。没有验证器怎么办不是每个站点都返回 ETag 或 Last-Modified。此时仍可利用内容指纹避免重复写库但省不下下载流量。不要自造If-None-Match客户端计算的正文哈希只有客户端知道服务器无法用它判断。也不要用“昨天请求过就跳过”替代验证这会把采集频率策略误当内容一致性证明。如果服务端错误地复用 ETag定期强制无条件请求是现实补救频率根据业务变化速度和成本决定。增量系统没有“一次配置永久正确”它需要监控 200/304 比例、发现数和内容更新时间。一处容易忽略的租户问题验证器与租户一起查询。tenant-a 有某 URL 的 ETag不代表 tenant-b 已经保存对应正文直接跨租户共用会让 tenant-b 首次请求收到 304却没有本地数据。共享缓存需要另外设计访问权限和引用计数不能为了省流量破坏隔离。本篇完整采集流水线请顺着run阅读列表页队列、详情 URL 去重、读取 validators、受限并发下载、304 分支、快照与 upsert、单条失败 rollback。每个计数字段都有明确增加位置。把下载、快照、解析和幂等写入组织成一次可报告的采集运行。fromdataclassesimportdataclass,fieldfromtypingimportProtocolfrom.concurrencyimportbounded_mapfrom.http_clientimportFetchResultfrom.parsingimportparse_detail,parse_listingfrom.repositoryimportJobRepositoryfrom.snapshotsimportFileSnapshotStoreclassFetcher(Protocol):asyncdeffetch(self,url:str,*,etag:str|NoneNone,last_modified:str|NoneNone,)-FetchResult:...dataclass(frozenTrue)classCrawlFailure:url:strmessage:strdataclassclassCrawlReport:list_pages:int0discovered:int0created:int0updated:int0unchanged:int0not_modified:int0failed:int0errors:list[CrawlFailure]field(default_factorylist)classCrawler:一次任务的应用服务单个详情失败不会抹掉其他成功结果。def__init__(self,fetcher:Fetcher,repository:JobRepository,snapshots:FileSnapshotStore,*,detail_concurrency:int4,)-None:self._fetcherfetcher self._repositoryrepository self._snapshotssnapshotsifdetail_concurrency1:raiseValueError(detail_concurrency must be positive)self._detail_concurrencydetail_concurrencyasyncdefrun(self,seed_url:str,*,tenant_id:str,max_pages:int10)-CrawlReport:ifmax_pages1:raiseValueError(max_pages must be at least 1)reportCrawlReport()pending[seed_url]seen_list_pages:set[str]set()detail_urls:list[str][]seen_details:set[str]set()whilependingandreport.list_pagesmax_pages:page_urlpending.pop(0)ifpage_urlinseen_list_pages:continueseen_list_pages.add(page_url)try:responseawaitself._fetcher.fetch(page_url)self._snapshots.save(response.url,response.body,response.headers)listingparse_listing(response.text,response.url)report.list_pages1fordetail_urlinlisting.detail_urls:ifdetail_urlnotinseen_details:seen_details.add(detail_url)detail_urls.append(detail_url)iflisting.next_urlandlisting.next_urlnotinseen_list_pages:pending.append(listing.next_url)exceptExceptionasexc:report.failed1report.errors.append(CrawlFailure(page_url,str(exc)))report.discoveredlen(detail_urls)requests:list[tuple[str,str|None,str|None]][]fordetail_urlindetail_urls:etag,last_modifiedself._repository.get_validators(tenant_id,detail_url)requests.append((detail_url,etag,last_modified))asyncdefdownload(request:tuple[str,str|None,str|None])-tuple[str,FetchResult|None,Exception|None]:detail_url,etag,last_modifiedrequesttry:responseawaitself._fetcher.fetch(detail_url,etagetag,last_modifiedlast_modified)returndetail_url,response,NoneexceptExceptionasexc:returndetail_url,None,exc downloadedawaitbounded_map(requests,download,limitself._detail_concurrency)fordetail_url,response,download_errorindownloaded:ifdownload_errorisnotNone:report.failed1report.errors.append(CrawlFailure(detail_url,str(download_error)))continueassertresponseisnotNonetry:ifresponse.not_modified:report.not_modified1continuesnapshotself._snapshots.save(response.url,response.body,response.headers)itemparse_detail(response.text,response.url)resultself._repository.upsert(tenant_id,item,etagresponse.etag,last_modifiedresponse.last_modified,snapshot_idsnapshot.snapshot_id,)self._repository.commit()setattr(report,result.action,getattr(report,result.action)1)exceptExceptionasexc:self._repository.rollback()report.failed1report.errors.append(CrawlFailure(detail_url,str(exc)))returnreport本篇课后练习用文字和请求头写出首次 200、第二次 304、内容变化后 200 的完整往返。解释为什么收到 304 时不能生成空快照也不能更新内容指纹。假设 1000 个详情中 950 个返回 304平均正文 50 KB忽略头部开销计算本次少传输了多少正文。下一篇将解决剩余 50 个详情怎样安全并发下载。
网站建设高端定制企业官网