Python中强大的蜘蛛(Web Crawler)系统.
A Powerful Spider(Web Crawler) System in Python.
Tutorial: http://docs.pyspider.org/en/latest/tutorial/
Documentation: http://docs.pyspider.org/
Release notes: https://github.com/binux/pyspider/releases
from pyspider.libs.base_handler import *
class Handler(BaseHandler):
crawl_config = {
}
@every(minutes=24 * 60)
def on_start(self):
self.crawl('http://scrapy.org/', callback=self.index_page)
@config(age=10 * 24 * 60 * 60)
def index_page(self, response):
for each in response.doc('a[href^="http"]').items():
self.crawl(each.attr.href, callback=self.detail_page)
def detail_page(self, response):
return {
"url": response.url,
"title": response.doc('title').text(),
}
pip install pyspiderpyspider, visit http://localhost:5000/WARNING: WebUI is open to the public by default, it can be used to execute any command which may harm your system. Please use it in an internal network or enable need-auth for webui.
Quickstart: http://docs.pyspider.org/en/latest/Quickstart/
Licensed under the Apache License, Version 2.0
Pyspider 命令行出错:语法不正确
执行 pyspider all 命令后,没有报错,但无法打开 web 页面
无法 pickle <function cli at 0x105140a60>: 它不是 pyspider.run.cli 的同一个对象
对非恒定时间比较的潜在侧通道攻击
输入 pyspider 时发生 PicklingError
无法运行 pyspider!
无法启动 pyspider - TypeError: 无法实例化抽象类 ScriptProvider,该类具有抽象方法 get_resource_inst
语法错误
webui.debug-itag height 是否增加了高度,不要让 iframe 太小。
希望调试页面 .iframe-box > iframe 添加 height=100%