小编Ada*_*m F的帖子

如何使用Python Scrapy模块列出我网站上的所有URL?

我想使用Python Scrapy模块从我的网站上抓取所有URL并将列表写入文件.我查看了示例,但没有看到任何简单的示例来执行此操作.

python web-crawler scrapy

20
推荐指数
2
解决办法
2万
查看次数

SQLAlchemy:order_by(None)用于joinload子条款查询?

我们在Python 2.7.7和Postgres 9.3上使用SQLAlchemy 0.9.8.

我们有一个查询使用joinedloads来使用单个查询完全填充一些Recipe对象.该查询创建一个大型SQL语句,执行时间为20秒 - 太长.这是Pastebin上呈现的SQL语句.

渲染的SQL有一个ORDER BY子句,Postgres解释说这是在这个查询上花费99%的时间的来源.这似乎来自ORM模型中的关系,它具有order_by子句.

但是,我们不关心为此查询返回结果的顺序 - 我们只关心查看单个对象时的顺序.如果我在呈现的SQL语句的末尾删除ORDER BY子句,则查询将在不到一秒的时间内执行 - 完美.

我们尝试在查询中使用.order_by(None),但这似乎没有任何效果.ORDER BY似乎与joinedloads有关,因为如果将joinedloads更改为lazyloads,它们就会消失.但我们需要加速加速.

如何让SQLAlchemy省略ORDER BY子句?


仅供参考,这是查询:

missing_recipes = cls.query(session).filter(Recipe.id.in_(missing_recipe_ids)) if missing_recipe_ids else []
Run Code Online (Sandbox Code Playgroud)

这是ORM类的摘录:

class Recipe(Base, TransactionalIdMixin, TableCacheMixin, TableCreatedModifiedMixin):
    __tablename__ = 'recipes'
      authors = relationship('RecipeAuthor', cascade=OrmCommonClass.OwnedChildCascadeOptions,
                           single_parent=True,
                           lazy='joined', order_by='RecipeAuthor.order', backref='recipe')
    scanned_photos = relationship(ScannedPhoto, backref='recipe', order_by="ScannedPhoto.position")
    utensils = relationship(CookingUtensil, secondary=lambda: recipe_cooking_utensils_table)
    utensil_labels = association_proxy('utensils', 'name')
Run Code Online (Sandbox Code Playgroud)

我们的query()方法看起来像这样(省略了一些joinloads):

@classmethod
def query(cls, session):
    query = query.options(
        joinedload(cls.ingredients).joinedload(RecipeIngredient.ingredient),
        joinedload(cls.instructions),
        joinedload(cls.scanned_photos),
        joinedload(cls.tags),
        joinedload(cls.authors),
    )
Run Code Online (Sandbox Code Playgroud)

python postgresql sqlalchemy

4
推荐指数
1
解决办法
1422
查看次数

标签 统计

python ×2

postgresql ×1

scrapy ×1

sqlalchemy ×1

web-crawler ×1