在python中处理大型文本文件

Ali*_*air 4 python text-files

我有一个非常大的文件(3.8G),它是我学校系统中用户的摘录.我需要重新处理该文件,以便它只包含他们的ID和电子邮件地址,逗号分隔.

我对此很少有经验,并希望将其用作Python的学习练习.

该文件的条目如下所示:

dn: uid=123456789012345,ou=Students,o=system.edu,o=system
LoginId: 0099886
mail: fflintstone@system.edu

dn: uid=543210987654321,ou=Students,o=system.edu,o=system
LoginId: 0083156
mail: brubble@system.edu
Run Code Online (Sandbox Code Playgroud)

我想获得一个看起来像这样的文件:

0099886,fflintstone@system.edu
0083156,brubble@system.edu
Run Code Online (Sandbox Code Playgroud)

任何提示或代码?

ig0*_*774 10

这对我来说实际上看起来像一个LDIF文件.该蟒蛇,LDAP库有一个纯Python LDIF处理库,可以帮助您的文件是否拥有一些可能在LDIF,如Base64编码值,进入折页等讨厌的陷阱

您可以像这样使用它:

import csv
import ldif

class ParseRecords(ldif.LDIFParser):
   def __init__(self, csv_writer):
       self.csv_writer = csv_writer
   def handle(self, dn, entry):
       self.csv_writer.writerow([entry['LoginId'], entry['mail']])

with open('/path/to/large_file') as input, with open('output_file', 'wb') as output:
    csv_writer = csv.writer(output)
    csv_writer.writerow(['LoginId', 'Mail'])
    ParseRecords(input, csv_writer).parse()
Run Code Online (Sandbox Code Playgroud)

编辑

因此,要从实时LDAP目录中提取,使用python-ldap库,您可能希望执行以下操作:

import csv
import ldap

con = ldap.initialize('ldap://server.fqdn.system.edu')
# if you're LDAP directory requires authentication
# con.bind_s(username, password)

try:
    with open('output_file', 'wb') as output:
        csv_writer = csv.writer(output)
        csv_writer.writerow(['LoginId', 'Mail'])

        for dn, attrs in con.search_s('ou=Students,o=system.edu,o=system', ldap.SCOPE_SUBTREE, attrlist = ['LoginId','mail']:
            csv_writer.writerow([attrs['LoginId'], attrs['mail']])
finally:
    # even if you don't have credentials, it's usually good to unbind
    con.unbind_s()
Run Code Online (Sandbox Code Playgroud)

阅读ldap模块的文档可能是值得的,尤其是示例.

请注意,在上面的示例中,我完全跳过提供过滤器,您可能希望在生产中执行此过滤器.LDAP中的过滤器类似于WHERESQL语句中的子句; 它限制返回的对象.微软实际上有一个很好的LDAP过滤器指南.LDAP过滤器的规范参考是RFC 4515.

类似地,如果在应用适当的过滤器之后可能存在数千个条目,您可能需要查看LDAP分页控件,尽管使用它会再次使示例更复杂.希望这足以让你开始,但如果有任何问题,请随时提出或打开一个新问题.

祝好运.


Sea*_*ira 5

假设每个条目的结构总是相同的,只需执行以下操作:

import csv

# Open the file
f = open("/path/to/large.file", "r")
# Create an output file
output_file = open("/desired/path/to/final/file", "w")

# Use the CSV module to make use of existing functionality.
final_file = csv.writer(output_file)

# Write the header row - can be skipped if headers not needed.
final_file.writerow(["LoginID","EmailAddress"])

# Set up our temporary cache for a user
current_user = []

# Iterate over the large file
# Note that we are avoiding loading the entire file into memory
for line in f:
    if line.startswith("LoginID"):
        current_user.append(line[9:].strip())
    # If more information is desired, simply add it to the conditions here
    # (additional elif's should do)
    # and add it to the current user.

    elif line.startswith("mail"):
        current_user.append(line[6:].strip())
        # Once you know you have reached the end of a user entry
        # write the row to the final file
        # and clear your temporary list.
        final_file.writerow(current_user)
        current_user = []

    # Skip lines that aren't interesting.
    else:
        continue
Run Code Online (Sandbox Code Playgroud)