我们假设我在python pandas中有下表
friend_description friend_definition
James is dumb dumb dude
Jacob is smart smart guy
Jane is pretty she looks pretty
Susan is rich she is rich
Run Code Online (Sandbox Code Playgroud)
在这里,在第一行中,'dumb'一词包含在两列中.在第二行中,'smart'包含在两列中.在第三行中,'pretty'包含在两列中,在最后一行中,'is'和'rich'包含在两列中.我想创建以下列:
friend_description friend_definition word_overlap overlap_count
James is dumb dumb dude dumb 1
Jacob is smart smart guy smart 1
Jane is pretty she looks pretty pretty 1
Susan is rich she is rich is rich 2
Run Code Online (Sandbox Code Playgroud)
我可以使用for循环来手动定义带有这些东西的新列,但我想知道pandas中是否有一个函数可以使这种类型的操作更加平滑.
处理此类字符串时,简单列表理解似乎是最快的方法:
\n\nIn [112]: df['word_overlap'] = [set(x[0].split()) & set(x[1].split()) for x in df.values]\n\nIn [113]: df['overlap_count'] = df['word_overlap'].str.len()\n\nIn [114]: df\nOut[114]:\n friend_description friend_definition word_overlap overlap_count\n0 James is dumb dumb dude {dumb} 1\n1 Jacob is smart smart guy {smart} 1\n2 Jane is pretty she looks pretty {pretty} 1\n3 Susan is rich she is rich {rich, is} 2\nRun Code Online (Sandbox Code Playgroud)\n\n单身的apply(..., axis=1):
In [85]: df['word_overlap'] = df.apply(lambda r: set(r['friend_description'].split()) &\n ...: set(r['friend_definition'].split()),\n ...: axis=1)\n ...:\n\nIn [86]: df['overlap_count'] = df['word_overlap'].str.len()\n\nIn [87]: df\nOut[87]:\n friend_description friend_definition word_overlap overlap_count\n0 James is dumb dumb dude {dumb} 1\n1 Jacob is smart smart guy {smart} 1\n2 Jane is pretty she looks pretty {pretty} 1\n3 Susan is rich she is rich {rich, is} 2\nRun Code Online (Sandbox Code Playgroud)\n\napply().apply(..., axis=1)方法:
In [23]: df['word_overlap'] = (df.apply(lambda x: x.str.split(expand=False))\n ...: .apply(lambda r: set(r['friend_description']) & set(r['friend_definition']),\n ...: axis=1))\n ...:\n\nIn [24]: df['overlap_count'] = df['word_overlap'].str.len()\n\nIn [25]: df\nOut[25]:\n friend_description friend_definition word_overlap overlap_count\n0 James is dumb dumb dude {dumb} 1\n1 Jacob is smart smart guy {smart} 1\n2 Jane is pretty she looks pretty {pretty} 1\n3 Susan is rich she is rich {is, rich} 2\nRun Code Online (Sandbox Code Playgroud)\n\n针对 40.000 行 DF 的计时:
\n\nIn [104]: df = pd.concat([df] * 10**4, ignore_index=True)\n\nIn [105]: df.shape\nOut[105]: (40000, 2)\n\nIn [106]: %timeit [set(x[0].split()) & set(x[1].split()) for x in df.values]\n223 ms \xc2\xb1 19.4 ms per loop (mean \xc2\xb1 std. dev. of 7 runs, 1 loop each)\n\nIn [107]: %timeit df.apply(lambda r: set(r['friend_description'].split()) & set(r['friend_definition'].split()), axis=1)\n3.65 s \xc2\xb1 46.6 ms per loop (mean \xc2\xb1 std. dev. of 7 runs, 1 loop each)\n\nIn [108]: %timeit df.apply(lambda x: x.str.split(expand=False)).apply(lambda r: set(r['friend_description']) & set(r['friend_definition']),\n ...: axis=1)\n4.63 s \xc2\xb1 84.7 ms per loop (mean \xc2\xb1 std. dev. of 7 runs, 1 loop each)\nRun Code Online (Sandbox Code Playgroud)\n
| 归档时间: |
|
| 查看次数: |
440 次 |
| 最近记录: |