lovemm 发表于 2022-8-2 16:32

使用Python爬虫一键爬取XX阁小说

本帖最后由 lovemm 于 2022-8-2 16:33 编辑

超级简单爬取笔趣阁小说的Python代码,只需要一个Python环境就能运行技术栈:requests,xpath直接上代码import os

import requests
from lxml import etree

def download_txt(name):
    params = {
      "keyword": name
    }
    host = "https://www.1biqug.com"
    resp = requests.get("https://www.1biqug.com/searchbook.php", params=params)
    html = resp.content.decode()
    html = etree.HTML(html)
    ret_list = html.xpath("//li/span[@class='s2']/a/@href")
    detail_url = host + ret_list
    resp = requests.get(detail_url)
    html = etree.HTML(resp.content.decode())
    ret_list = html.xpath("//div[@id='list']//dd//a/@href")
    print(ret_list)
    if not os.path.exists("./{}".format(name)):
      os.mkdir("./{}".format(name))
    for ret in ret_list:
      url = host + ret
      resp = requests.get(url)
      info = resp.content.decode()
      html = etree.HTML(info)
      title = html.xpath("//h1/text()")
      print(title)
      path = os.path.join(name, title + ".html")
      path = path.replace("*", "")
      with open(path, 'w', encoding="utf8") as f:
            f.write(info)
    print(name, "下载完成了")

if __name__ == '__main__':
    story = input("请输入小说名")
    download_txt(story)




花香雨润 发表于 2022-8-2 16:32

谢谢分享!

尹志平 发表于 2022-8-2 16:33

多谢楼主分享

去污存清 发表于 2022-8-2 16:38

谢谢分享

笑东风丶 发表于 2022-8-2 16:55

感谢大佬分享,大佬666666

摩擦摩擦 发表于 2022-8-2 17:13

6666666666

里面那边 发表于 2022-8-2 17:14

谢谢大牛

xt686 发表于 2022-8-2 17:16


感谢楼主的分享

sfcdnlt 发表于 2022-8-2 17:20

谢谢大佬分享

qq994589328 发表于 2022-8-2 21:28

先收藏了,万一用到呢
页: [1] 2
查看完整版本: 使用Python爬虫一键爬取XX阁小说