万一的爸爸 发表于 2022-9-14 03:48

感谢大佬分享,大佬辛苦

ymx0312666 发表于 2022-9-14 14:26

66666666

八度幸福 发表于 2022-9-18 00:25

是进口的吗?

在你心里 发表于 2022-9-18 10:18

哦哦哦哦

lyxje 发表于 2022-9-18 10:31

666666666666666

xiaoxiongii 发表于 2022-9-18 10:59

正需要,支持楼主,在大牛我只看好你!

飞天柚子 发表于 2022-9-18 11:08

感谢分享

haof1 发表于 2022-9-20 00:45

import os
import re
import time
from urllib import request
from bs4 import BeautifulSoup


def get_last_page(text):
    return int(re.findall('[^/$]\d*', re.split('/', text)[-1]))


def html_parse(url, headers):
    time.sleep(3)
    resp = request.Request(url=url, headers=headers)
    res = request.urlopen(resp)
    html = res.read().decode("utf-8")
    soup = BeautifulSoup(html, "html.parser", from_encoding="utf-8")
    return soup


headers = {
    'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36 Edg/91.0.864.59'
}

url = "https://www.2meinv.com/"
for p in range(1, 10 + 1):
    next_url = url + "index-" + str(p) + ".html"
    soup = html_parse(next_url, headers)
    link_node = soup.findAll('div', attrs={"class": "dl-name"})
    for a in link_node:
      path = "G:/spider/image/2meinv/"
      href = a.find('a', attrs={'target': '_blank'}).get('href')
      no = re.findall('[^-$][\d]', href) + re.findall('[^-$][\d]', href)
      first_url = url + "/article-" + no + ".html"
      title = a.find('a', attrs={'target': '_blank'}).text
      path = path + title + "/"
      soup = html_parse(href, headers)
      count = soup.find('div', attrs={'class': 'des'}).find('h1').text
      last_page = get_last_page(count)
      for i in range(1, last_page + 1):
            next_url = url + "/article-" + no + "-" + str(i) + ".html"
            soup = html_parse(next_url, headers)
            image_url = soup.find('img')['src']
            image_name = image_url.split("/")[-1]
            fileName = path + image_name
            if not os.path.exists(path):
                os.makedirs(path)
            if os.path.exists(fileName):
                continue
            request.urlretrieve(image_url, filename=fileName)
            request.urlcleanup()
      print(title, "下载完成了")

z1z1z 发表于 2022-9-20 10:30

6666666666666

a995799499 发表于 2022-9-20 10:35

谢谢分享
页: 1 2 3 4 [5] 6 7 8 9 10 11 12 13 14
查看完整版本: 【TV/盒子】某猫最新版破了,免广告极速秒播,文末赠苹果(IOS)端免费观影追剧软...