宜配屋

使用python BeautifulSoup库抓取58手机维修信息

yipeiwu_com6年前 (2020-03-06)Python爬虫

直接上代码：

复制代码代码如下:

#!/usr/bin/python
# -*- coding: utf-8 -*-

import urllib

import os,datetime,string

import sys

from bs4 import BeautifulSoup

reload(sys)

sys.setdefaultencoding('utf-8')

__BASEURL__ = 'http://bj.58.com/'

__INITURL__ = "http://bj.58.com/shoujiweixiu/"

soup = BeautifulSoup(urllib.urlopen(__INITURL__))

lvlELements = soup.html.body.find('div','selectbarTable').find('tr').find_next_sibling('tr')('a',href=True)

f = open('data1.txt','a')

for element in lvlELements[1:]:

f.write((element.get_text()+'\n\r' ))

url = __BASEURL__ + element.get('href')

print url

soup = BeautifulSoup(urllib.urlopen(url))

lv2ELements = soup.html.body.find('table','tblist').find_all('tr')

    for item in lv2ELements:
        addr = item.find('td','t').find('a').get_text()
        phone = item.find('td','tdl').find('b','tele').get_text()
        f.write('地址：'+addr +' 电话:'+ phone + '\r\n\r')

f.close()

直接执行后，存在 data1.txt中就会有商家的地址和电话等信息。
BeautifulSoup api 的地址为： http://www.crummy.com/software/BeautifulSoup/bs4/doc/

使用python BeautifulSoup库抓取58手机维修信息

相关文章

python模拟新浪微博登陆功能(新浪微博爬虫)

python协程gevent案例爬取斗鱼图片过程解析

Python爬虫实现验证码登录代码实例

一步步教你用python的scrapy编写一个爬虫

Python3爬虫学习之爬虫利器Beautiful Soup用法分析

© YiPeiWu.com 【宜配屋】粤ICP备17031333号

Powered By Z-BlogPHP. Theme by TOYEAN.

宜配屋

使用python BeautifulSoup库抓取58手机维修信息

相关文章

python模拟新浪微博登陆功能(新浪微博爬虫)

python协程gevent案例 爬取斗鱼图片过程解析

Python爬虫实现验证码登录代码实例

一步步教你用python的scrapy编写一个爬虫

Python3爬虫学习之爬虫利器Beautiful Soup用法分析

© YiPeiWu.com 【宜配屋】 粤ICP备17031333号 var _hmt = _hmt || [];(function() { var hm = document.createElement("script"); hm.src = "https://hm.baidu.com/hm.js?8aa60ae04b767b2af31903508928acc0"; var s = document.getElementsByTagName("script")[0]; s.parentNode.insertBefore(hm, s);})();

Powered By Z-BlogPHP. Theme by TOYEAN.

python协程gevent案例爬取斗鱼图片过程解析

© YiPeiWu.com 【宜配屋】粤ICP备17031333号