我想使用Python Scrapy模块从我的网站上抓取所有URL并将列表写入文件.我查看了示例,但没有看到任何简单的示例来执行此操作.
我们在Python 2.7.7和Postgres 9.3上使用SQLAlchemy 0.9.8.
我们有一个查询使用joinedloads来使用单个查询完全填充一些Recipe对象.该查询创建一个大型SQL语句,执行时间为20秒 - 太长.这是Pastebin上呈现的SQL语句.
渲染的SQL有一个ORDER BY子句,Postgres解释说这是在这个查询上花费99%的时间的来源.这似乎来自ORM模型中的关系,它具有order_by子句.
但是,我们不关心为此查询返回结果的顺序 - 我们只关心查看单个对象时的顺序.如果我在呈现的SQL语句的末尾删除ORDER BY子句,则查询将在不到一秒的时间内执行 - 完美.
我们尝试在查询中使用.order_by(None),但这似乎没有任何效果.ORDER BY似乎与joinedloads有关,因为如果将joinedloads更改为lazyloads,它们就会消失.但我们需要加速加速.
如何让SQLAlchemy省略ORDER BY子句?
仅供参考,这是查询:
missing_recipes = cls.query(session).filter(Recipe.id.in_(missing_recipe_ids)) if missing_recipe_ids else []
Run Code Online (Sandbox Code Playgroud)
这是ORM类的摘录:
class Recipe(Base, TransactionalIdMixin, TableCacheMixin, TableCreatedModifiedMixin):
__tablename__ = 'recipes'
authors = relationship('RecipeAuthor', cascade=OrmCommonClass.OwnedChildCascadeOptions,
single_parent=True,
lazy='joined', order_by='RecipeAuthor.order', backref='recipe')
scanned_photos = relationship(ScannedPhoto, backref='recipe', order_by="ScannedPhoto.position")
utensils = relationship(CookingUtensil, secondary=lambda: recipe_cooking_utensils_table)
utensil_labels = association_proxy('utensils', 'name')
Run Code Online (Sandbox Code Playgroud)
我们的query()方法看起来像这样(省略了一些joinloads):
@classmethod
def query(cls, session):
query = query.options(
joinedload(cls.ingredients).joinedload(RecipeIngredient.ingredient),
joinedload(cls.instructions),
joinedload(cls.scanned_photos),
joinedload(cls.tags),
joinedload(cls.authors),
)
Run Code Online (Sandbox Code Playgroud)